Back to Blog

How Do ChatGPT Detectors Work? Explained Simply (2026)

By Arsalan Amin July 12, 2026 Updated September 5, 2026 7 min read
How Do ChatGPT Detectors Work? Explained Simply (2026)

You paste an essay, a progress bar spins, and a tool declares '87% AI-generated.' It feels like magic or fraud, depending on which side of the flag you're standing on. It's neither. Under the hood, nearly every ChatGPT detector runs on two measurable signals, and once you understand them, the whole detection world stops being mysterious.

Signal one: perplexity

Language models write by predicting the next word. When ChatGPT writes, it usually picks a highly probable word, then another, then another. The result is text a model finds easy to predict. That predictability has a name: low perplexity.

Monochrome abstract wave pattern creating a hypnotic design with repetition and depth.

Detectors run your text through their own model and ask, at each word, how surprised the model is. Consistently unsurprised means machine-shaped. Human writing surprises constantly, because people pick words for meaning, sound, and habit rather than probability.

Signal two: burstiness

Humans write in bursts. Three long sentences, then a fragment. A rambling aside, then a verdict. AI output keeps a steady cadence, sentence after sentence of similar length and structure. Detectors measure that variance directly. Low variance reads as machine, high variance reads as human.

The third ingredient: trained classifiers

Modern detectors layer a trained model on top: feed it millions of known-AI and known-human samples and let it learn the difference. This makes them better at catching lightly edited AI text, and it's also why they inherit bias. Whatever looked machine-like in training data gets flagged in the wild, including formulaic human writing and non-native English.

That bias is measurable. Our breakdown of whether AI detectors are accurate covers the false-positive patterns in detail.

Why this explains everything else

Diagram of how a ChatGPT detector reaches a number, using the two signals this post describes plus the trained layer on top. Signal one is perplexity: models write by picking a highly probable next word, then another, so the result is text a model finds easy to predict, and detectors run your text through their own model and ask at each word how surprised it is, where consistently unsurprised reads as machine-shaped. Signal two is burstiness: humans write in bursts, three long sentences then a fragment, a rambling aside then a verdict, while AI output keeps a steady cadence of similar sentence lengths, and detectors measure that variance directly, where low variance reads as machine and high variance reads as human. The third ingredient is a trained classifier layered on top, fed millions of known-AI and known-human samples, which is better at catching lightly edited AI text and is also the reason detectors inherit bias, since whatever looked machine-like in training gets flagged in the wild, including formulaic human writing and non-native English. Four outcomes follow from the two signals: raw ChatGPT flags because it is maximum predictability and minimum burstiness, the easiest possible target; synonym-swapping fails because word substitutions barely move rhythm and rhythm is half the score, leaving one dial untouched; polished human writing flags because grammar tools smooth out the very human variance the second signal looks for; and real rewriting works because it rebuilds both signals at once, restoring the variation detectors expect from people. The footer restates the post's point that the detector never watched the writing happen, so every verdict is an inference from statistics rather than proof of authorship.
  • Why raw ChatGPT flags: maximum predictability, minimum burstiness. The easiest possible target.
  • Why synonym-swapping fails: word substitutions barely move rhythm, and rhythm is half the score.
  • Why polished human writing gets flagged: grammar tools smooth out the human variance.

And why proper rewriting works: a real AI humanizer rebuilds both signals at once, restoring the variation detectors expect from people.

You can watch these signals move yourself: paste any text into a free AI detector, edit, and re-check. Ten minutes of that teaches more than any article.

Frequently asked questions

Do ChatGPT detectors read metadata or hidden watermarks?

The tools in common use score the text itself. Watermarking research exists at the model-vendor level, but today's detection verdicts come from statistics, not hidden tags.

Can a detector tell which AI wrote something?

Not reliably. The statistical fingerprint is similar across models, which is why reports say 'AI-generated' rather than naming ChatGPT.

Why do two detectors disagree on the same text?

Different models, different training data, different thresholds. Disagreement is normal, which is worth remembering when a single score gets treated as proof.

Running the model on your own machine changes nothing

A common assumption is that a model running locally, in LM Studio or something similar, with no API call ever leaving the machine, must produce output that is somehow invisible. That is not how any of this works. Detection happens on the text itself, after the fact, and the statistical fingerprint comes from how large language models generate text in general rather than from which company hosts the model. A 7B open weight model on a laptop and a frontier cloud model share the same pull toward smooth, average, high probability phrasing, because that is what makes a language model coherent.

If anything a local draft needs more attention, not less, since smaller models vary more in quality and are more prone to small factual slips worth catching early. Beyond that the two signals above behave identically: vary sentence rhythm deliberately, replace default phrasing with words you would actually use, break the model's even paragraphing, and add a detail it could not have known. Reading the result aloud is a decent proxy for both signals at once, because people and detectors are responding to the same thing, natural variation.

Switching models is not evasion either

The same logic settles the question people ask about DeepSeek, or Claude, or whatever launched last month. Detectors are not matching output against a known fingerprint for one vendor's model, so a detector does not need to have seen a given model to flag what it produces. It is scoring the general statistical pattern of machine generated prose. DeepSeek's architecture, training data and research team are genuinely different from OpenAI's, and the end product still leans toward the same predictable, evenly paced writing, because that tendency is what coherence costs.

Phrasing habits do differ slightly between model families, in the way any two writers differ, but not in a way that changes whether the text gets caught. Assuming a newer or less mainstream model flies under the radar is the same mistake as assuming a locally run one does, and it is why detector reports say AI generated rather than naming a product. Who built the model and where it runs change nothing about how the output reads.

What is perplexity actually measuring, and how is it calculated?

Perplexity is a number derived from probability, and the calculation runs on tokens, not on words. A language model never sees letters or words. It splits input into tokens, which are subword chunks produced by byte pair encoding. OpenAI's tiktoken documentation puts the average at roughly four bytes per token and gives the example of the word encoding being split into the pieces encod and ing. The model then assigns a probability to every possible next token at every position in the sequence. A detector pushes your text through a model, reads off the probability the model assigned to the token that actually appeared, and averages that across the passage. Consistently high assigned probabilities mean low perplexity, which sits at the machine-shaped end of the scale.

The research literature has a sharper version of the same idea. DetectGPT, published by Mitchell and colleagues in 2023, measures the shape of the probability surface around your text rather than probability alone, on the observation that machine-written text tends to occupy negative curvature regions of a model's log probability function. On fake news generated by a 20 billion parameter model it reported 0.95 AUROC against 0.81 for the strongest zero-shot baseline it was compared with. The consequence for anyone reading a score is that perplexity is not one fixed quantity. Each tool computes it against a different reference model, which is a large part of why two detectors return different numbers on identical text.

Why does burstiness need a long passage before it means anything?

Burstiness is variance, and variance needs samples. In a six-sentence paragraph a detector is estimating the spread of sentence lengths from six data points, so one unusual sentence drags the estimate a long way. At three hundred words the same estimate rests on roughly twenty sentences and stops swinging on single edits. This is not a theory about how detectors ought to behave, it is what the vendors document. Turnitin raised its minimum from 150 words to 300 in May 2023 and said directly that accuracy increases with more text. ZeroGPT's FAQ recommends at least 150 to 200 words and says 500 to 1,000 or more improves stability. GPTZero states that its document-level accuracy is higher than its paragraph-level accuracy, which is higher than its sentence-level accuracy. Three companies, three admissions that short text breaks the method.

Why is the answer always a probability instead of a yes or no?

Because the detector never observed the writing happening. It sees only finished text, and finished text carries no proof of its own origin, so every verdict is an inference from statistics. That is why each tool wraps its output in confidence language rather than a verdict. GPTZero's API returns a document classification of HUMAN_ONLY, MIXED or AI_ONLY alongside a confidence category of high, medium or low, and the company states that error rates fall below one percent only inside the high confidence band. Turnitin reports its sentence-level false positive rate at around 4 percent and adds that 54 percent of the time a wrongly highlighted sentence sits directly next to genuine AI writing, which is a classifier smearing a boundary rather than identifying an author.

There is also a ceiling no amount of engineering removes. Sadasivan and colleagues published a framework in 2023 linking the performance of the best possible detector to the statistical distance between human and machine text distributions. As models get better at writing like people, that distance shrinks and the ceiling comes down with it. The same paper showed recursive paraphrasing substantially reduces detection rates while degrading text quality only slightly. That is the structural reason a detector says 87 percent rather than yes.

What proves detectors are scoring predictability rather than authorship?

A single experiment settles it. Liang and colleagues at Stanford took 88 US eighth-grade essays, all genuinely written by children, and asked ChatGPT to simplify the word choices as if written by a non-native speaker. The content, the argument and the human author were unchanged. Average misclassification as AI-generated across seven detectors rose from 5.19 percent to 56.65 percent. Running the reverse intervention on 91 human-written TOEFL essays, asking ChatGPT to enrich the word choices to sound more like a native speaker, dropped the average false positive rate from 61.22 percent to 11.77 percent. Nothing about who wrote those essays changed in either direction. Only the predictability of the vocabulary changed, and the verdicts moved by tens of percentage points. That is the clearest available demonstration that these tools score a statistical property of prose and then let the reader interpret it as authorship.

How this was made: this post predates this site's "How this was made" disclosure convention, which was added 2026-08-16. Drafting was AI-assisted with human editing. This paragraph and the diagram above were added during a 2026-09-05 accuracy pass; the diagram restates points the post already made, and no claim in this post was re-verified beyond what is stated above.

AI DetectionChatGPTHow It Works

Keep reading

Free trial available

Make AI writing sound naturally human.

Paste your draft, pick a tone, and get clear, natural writing in seconds, with a built-in AI detector to check your work.

Humanize My Text

No credit card required • Cancel anytime

Unlimited humanization5 writing stylesAI detector includedPrivate & secureInstant results