Back to Blog

Are AI Detectors Accurate? The Honest Numbers (2026)

By Arsalan Amin July 12, 2026 Updated September 7, 2026 10 min read
Are AI Detectors Accurate? The Honest Numbers (2026)

Every AI detector advertises a number north of 98%. Meanwhile, universities keep walking back detector-based accusations and Reddit fills up with students whose hand-written essays got flagged. Someone is wrong. The truth is more useful than either side admits.

What the accuracy claims actually measure

Vendors test on benchmark datasets: known-AI text on one side, known-human text on the other. On clean benchmarks, modern detectors genuinely do score in the high nineties. The problem is that almost nothing about real-world writing is clean. Real text gets grammar-polished, translated, dictated, edited by committee, or written by tired people following rigid formats. Every one of those factors drags human writing toward what a detector calls machine-like.

Woman using multiple screens for cybersecurity tasks in a cozy home office

Where detectors go wrong most

  • Non-native English writers. Safer vocabulary and simpler constructions read as predictable. Stanford researchers found some detectors flagged over half of essays by non-native speakers.
  • Formulaic formats. Lab reports, legal boilerplate, five-paragraph essays. Uniform by design, flagged for being uniform.
  • Grammar-tool polish. Heavy editing strips the natural variation that reads as human.
  • Short samples. Under a few hundred words, the statistics get thin and the verdicts get noisy.

The disagreement problem

Here's the test anyone can run: take one essay and feed it to three detectors. You will routinely get three different verdicts. GPTZero says 80% AI, another tool says 12%, a third says 45%. If detection were as accurate as the marketing says, the tools would agree. They don't, because each weighs perplexity, burstiness, and its own training data differently.

We've documented this per tool: see how accurate Turnitin's AI detection is and

the same deep-dive on GPTZero's accuracy. Same pattern in both: good at raw AI, unreliable at edges.

What this means for you

  • If you write honestly: keep drafts and version history. Process evidence beats percentages every time.

Check your work in a free AI detector before submitting anything that matters, so a false flag never ambushes you.

If you write with AI help: rewrite properly. A quality AI humanizer plus your own edits restores the variation detectors need to see.

Frequently asked questions

Which AI detector is most accurate?

Rankings shift with every model update, and the leaders disagree with each other on the same text. Whichever tool you face, the same preparation works: natural variation, real specifics, process evidence.

Can AI detectors be wrong?

Yes, in both directions. False positives on human writing are common enough that several universities disabled detector features entirely.

Should teachers trust AI detector scores?

As a conversation starter, not a verdict. Every major vendor, including Turnitin, says exactly this in its own documentation.

How to judge a detector's score, and the tools people ask about

Everything above is about whether detectors are accurate in general. This section is about the number actually in front of you. The most useful move is to stop asking which detector is right and start testing the detector itself, because a tool you have calibrated against text whose origin you already know is worth far more than a tool you have only read reviews of.

The three-passage test: calibrate the detector before you trust it

The disagreement test above compares tools against each other. This one is different: it tests a single tool against text whose author you already know. Take three passages. Something you wrote by hand months ago, something a model generated and you never touched, and something you wrote yourself and then edited heavily. Run all three through the detector. A tool worth using should separate the first two clearly. If it cannot, the number it gives your real work is noise, and no amount of confident marketing changes that.

Then run the second half of the test. Take one of those passages and feed it again as a 150-word excerpt. If the score moves substantially on the shorter sample, you have found that tool's length limit, and you now know not to trust any result it gives you below it. The whole thing takes about ten minutes and tells you more than any review, including this one.

What a detector percentage is not

A detector returns a percentage, and the useful question is not whether that percentage is right but what it is a percentage of. It is not the probability that a specific person did or did not write the passage. It is a statistical judgment about how predictable the writing is. Those are two different claims, and the gap between them causes most of the arguments about detector results.

That gap runs in both directions. A low score is weak supporting evidence at best and cannot prove you wrote something yourself, which is why drafts, version history, and notes remain stronger proof of authorship than any detector output. And when two tools disagree sharply, neither one is necessarily broken. They were trained on different corpora, use different thresholds, and calibrate against different model generations. Sharp disagreement usually means your text sits near a decision boundary, so the honest reading is that both numbers are inconclusive rather than that you get to pick the one you prefer.

The minimum length problem

Roughly 300 words is the practical floor for this class of tool. Below that there is not enough statistical signal, and scores swing hard on small edits, sometimes on a single changed sentence. Treat that 300 as a rule of thumb rather than a measured constant, because no vendor publishes a validated threshold and it will vary by tool. The point is that a floor exists and that short samples sit under it.

This gives you a clean rule for when a score is worth acting on at all. Worth acting on: a consistently high score across multiple detectors, on a passage longer than 300 words, where you already know the text started as model output and you are trying to find out whether your editing did enough. Not worth acting on: a single score on a short passage, a score on writing you produced yourself, or a score you are trying to optimise toward a target number. Optimising for a detector is a losing game, because detectors retune and the target moves.

Bundled detectors and specialist detectors have different incentives

Some detectors are standalone products. Others are one feature inside a general AI assistant that also does writing, summarising, browsing, and chat. That distinction matters more than any accuracy number, because it changes what the tool was optimised for. In a bundled product, detection competes for engineering attention against every other feature on the roadmap. A company whose entire business depends on detection accuracy has a different incentive to keep retuning against the latest models than a company maintaining a detector bolted onto a broader suite.

This is a structural argument, not an accusation about any specific product, and it is not automatically a problem. It is a trade-off worth naming before you rely on a score for anything that matters. In practice: a bundled detector is fine for a quick gut check on your own writing. For anything graded or high-stakes, cross-check against a second tool built by someone whose main product is detection.

The named tools, briefly and honestly

The first thing to say is what applies to all of them. Every mainstream detector, including Merlin, Reilaa, and Hastewire, scores the same two statistical properties: perplexity and burstiness. None of them reads a watermark or a hidden signature, and none of them can tell you which model produced a passage, which is a different question from how machine-like the statistics look. That means they all inherit the same blind spots on short text, technical and formulaic writing, and non-native English. No independent, systematically tested accuracy figure exists in public for Merlin, Reilaa, or Hastewire, and any review that quotes a confident percentage for a tool it never tested is guessing. Free tiers and pricing in this category change frequently, so check current terms directly rather than trusting any review.

On the better-known tools, the qualitative differences are real even where the numbers are not. GPTZero is widely adopted in education, offers sentence-level scoring and LMS integrations, and is openly weak on heavily edited or humanised content. Winston AI is aimed at publishers and SEO teams working with long-form content. Originality.ai is built for agency-scale bulk scanning and API use, and carries a well-known reputation for false positives on genuinely human writing. Copyleaks is API-first with multilingual coverage, which is why it shows up in custom integrations. Turnitin sits behind institutional access, so most students cannot run their own work through it before submitting, and that is exactly why a clean score from a consumer detector tells you very little about what an instructor will see.

One correction to something this site published previously. An earlier comparison here listed 2025 accuracy figures of about 83% for GPTZero, 89% for Winston AI, 91% for Originality.ai, 93% for Turnitin, and 95% for Copyleaks. Those numbers were published with no test, no methodology, and no source cited, so treat them as unverified claims rather than measured results. They also describe a moment that has passed, since detectors retune against new model generations continuously. We are leaving them here labelled rather than quietly deleting them, because the ranking they implied has been repeated elsewhere and it deserves the caveat more than it deserves the citation.

Which detector accuracy numbers can you actually check, and who published them?

Diagram contrasting the two sets of detector accuracy numbers this post already collects, neither of them fabricated, because they measure different tools on different text in different years. On the vendor side: Winston AI's homepage advertises 99.98 percent; ZeroGPT's FAQ says it is pushing toward above 98 percent on evaluations it calls internal; GPTZero's technology page claims 96.5 percent on mixed documents, a false positive rate no higher than 1 percent and a TOEFL false positive rate of 1.1 percent; and Turnitin states document-level false positives under 1 percent but only for documents where it detects 20 percent or more AI writing, with a sentence-level rate around 4 percent. Every one of those figures was produced by the company selling the product. On the independent side: Weber-Wulff and colleagues in the International Journal for Educational Integrity in 2023 tested 14 tools and, under their strictest scoring method, put Turnitin first at 76 percent, Go Winston at 67 percent, ZeroGPT at 59 percent and GPTZero at 54 percent, concluding that all tools scored below 80 percent and only 5 over 70; a 2025 test in PeerJ Computer Science on 72 journal abstracts put GPTZero at 97.22 percent accuracy with a 0.00 percent false positive rate and ZeroGPT at 64.35 percent accuracy with a 16.67 percent false positive rate. A third panel shows the number that moves depending on who wrote the text: Liang and colleagues ran 91 human-written TOEFL essays through seven detectors and recorded an average false positive rate of 61.22 percent, with all seven agreeing on 18 of the 91 and 89 of the 91 flagged by at least one tool, while the same detectors misclassified an average of 5.19 percent of 88 US eighth-grade essays, a more than tenfold difference driven by authorship rather than by the task. The final panel restates Vanderbilt's public arithmetic: its August 2023 statement notes the university submitted 75,000 papers to Turnitin in 2022, so a 1 percent false positive rate would have meant roughly 750 papers incorrectly labelled as containing AI writing, which is why it disabled the detector among other reasons. One percent is a good error rate for a statistical classifier and also hundreds of difficult conversations a year at one institution.

Two sets of numbers exist for every detector, and they do not agree. The companies publish their own figures, which sit in the high nineties. Independent researchers publish theirs, which mostly land between 50 and 80 percent. Neither set is fabricated; they measure different tools on different text in different years. Every figure below is attributed to a source you can open yourself.

What do the detector companies claim about their own accuracy?

Turnitin states a document-level false positive rate of less than 1 percent, but only for documents where it detects 20 percent or more AI writing, and puts its sentence-level rate at around 4 percent. GPTZero's technology page claims 96.5 percent accuracy on mixed documents containing both human and AI writing, a false positive rate of no more than 1 percent, and a TOEFL false positive rate of 1.1 percent. ZeroGPT's FAQ says it is pushing toward above 98 percent on internal evaluations, with the word internal doing real work. Winston AI's homepage advertises 99.98 percent. Every one of those numbers was produced by the company selling the product.

What did independent researchers measure on the same tools?

The broadest published test of consumer detectors is Weber-Wulff and colleagues in the International Journal for Educational Integrity in 2023, covering 14 tools including Turnitin. Under their strictest scoring method, Turnitin ranked first at 76 percent, Go Winston reached 67 percent, ZeroGPT scored 59 percent and GPTZero scored 54 percent. Their conclusion sentence is blunt: all tools scored below 80 percent of accuracy and only 5 over 70. A newer peer-reviewed test in PeerJ Computer Science in 2025, run on 72 journal abstracts, put GPTZero at 97.22 percent accuracy with a 0.00 percent false positive rate and ZeroGPT at 64.35 percent accuracy with a 16.67 percent false positive rate. Those two studies used different text two years apart against different model generations, which is precisely why no stable ranking of detectors exists.

How badly do detectors treat non-native English writers, in numbers?

The Stanford finding mentioned earlier has an exact figure attached. Liang and colleagues ran 91 human-written TOEFL essays through seven detectors, GPTZero and ZeroGPT among them, and recorded an average false positive rate of 61.22 percent. All seven agreed on 18 of the 91 essays, and 89 of the 91 were flagged by at least one tool. On 88 US eighth-grade essays the same detectors misclassified an average of 5.19 percent. Same tools, same task, a difference of more than ten times, driven by who wrote the text. GPTZero reports a 1.1 percent TOEFL false positive rate today after working on this since 2022, which is a vendor figure covering one tool and does not retire the finding for the category.

What happens when you multiply a one percent error rate by a real university?

Vanderbilt University did that arithmetic in public. Its August 2023 statement notes the university submitted 75,000 papers to Turnitin in 2022, so a 1 percent false positive rate would have meant roughly 750 student papers incorrectly labelled as containing AI writing. Vanderbilt disabled Turnitin's AI detector for that reason among others, and observes in the same statement that Turnitin publishes no detailed account of how the detection works. One percent is a good error rate for a statistical classifier. It is also hundreds of difficult conversations a year at one institution.

How this was made: this post predates this site's "How this was made" disclosure convention, which was added 2026-08-16. Drafting was AI-assisted with human editing. This paragraph and the diagram above were added during a 2026-09-07 accuracy pass; the diagram restates points the post already made, and no claim in this post was re-verified beyond what is stated above.

AI DetectionFalse PositivesAccuracy

Keep reading

Free trial available

Make AI writing sound naturally human.

Paste your draft, pick a tone, and get clear, natural writing in seconds, with a built-in AI detector to check your work.

Humanize My Text

No credit card required • Cancel anytime

Unlimited humanization5 writing stylesAI detector includedPrivate & secureInstant results