
Turnitin says its AI detector is highly accurate. Students who've been wrongly flagged say otherwise. Both can be true at once, and understanding why matters whether you write everything yourself, use AI to brainstorm, or humanize AI-assisted drafts. Here's what the accuracy numbers actually mean, where the detector goes wrong, and what to do about it.
What 'accuracy' actually means for an AI detector
Turnitin has claimed a false-positive rate of around 1% for documents at the sentence-and-above level. That sounds tiny until you scale it. Across millions of submissions each term, a 1% error rate means tens of thousands of genuinely human essays flagged as AI. Accuracy is also measured on clean lab data. Real student writing is messier, and that's exactly where error rates climb.

Why false positives happen
The detector scores predictability, not authorship. Anything that makes human writing statistically 'smooth' can trip it:
- Formulaic academic style. Five-paragraph structures and standard topic sentences look uniform by design.
- Non-native English writers. Simpler, safer phrasing reads as more predictable, and studies have shown these writers are flagged disproportionately.
- Technical and scientific writing. Fixed terminology and rigid structure lower perplexity naturally.
- Grammar tools. Heavy Grammarly-style polishing smooths out exactly the variation detectors expect from humans.
What to do if you're falsely flagged
- Don't panic and don't confess to something you didn't do. A score is not proof, and Turnitin itself tells institutions to treat it as an indicator rather than a verdict.
- Gather your evidence: version history in Google Docs or Word, drafts, notes, and browser history all show the work happening over time.
- Ask how the number was produced. Many instructors don't know the detector's error profile, and the 1%-per-document figure understates per-passage errors.
- Request a human review. Policies at most institutions require more than a detector score to sustain an accusation.
Protecting yourself before you submit
Whether your draft is fully yours or AI-assisted, check how it reads to a machine before Turnitin does: run it through a free AI detector first. If your genuinely human writing scores high, vary your sentence lengths and add specific personal detail. The same fixes that make writing better also make it read more human.
If you drafted with AI help, rewrite it properly rather than submitting raw output. Our guide on making AI-assisted writing read naturally walks through what actually changes a score.
The bottom line
Turnitin's AI detection is a useful statistical signal with a real error rate. It's accurate enough to catch lazy copy-paste AI, and unreliable enough that innocent writers get flagged every term. Know how it works, keep evidence of your process, and check your own text before you submit. That preparation costs minutes and can save a semester.
Frequently asked questions
Can Turnitin's AI detector be wrong?
Yes. Turnitin acknowledges false positives, and independent testing has found error rates above the marketing numbers, especially for non-native English writers and formulaic styles.
What score counts as 'AI-written'?
There's no official cutoff. Turnitin shows instructors a percentage of the document it believes is AI-generated, and each institution decides how to act. Many ignore scores below 20% entirely.
Do professors see the AI score automatically?
If their institution has the AI-writing feature enabled, yes. It appears alongside the similarity report. Some institutions have disabled it precisely because of false positives.
What actually happened at Vanderbilt, and does the 750 figure mean 750 students were falsely accused?
On 16 August 2023 Vanderbilt University published a post on its Brightspace blog titled Guidance on AI Detection and Why We're Disabling Turnitin's AI Detector. The number everybody quotes from it is almost always quoted wrong. Vanderbilt wrote that it submitted 75,000 papers to Turnitin in 2022, and that at Turnitin's claimed 1 percent false positive rate, around 750 student papers could have been incorrectly labelled as having some of it written by AI. That is arithmetic performed on a vendor's own claim, not a tally of students who were accused. Turnitin's AI detector did not exist in 2022, so it could not have produced 750 real flags that year. If you see the figure described as 750 confirmed false positives, whoever wrote that did not read the source.
What Vanderbilt decided is real enough. It disabled the tool for the foreseeable future, citing that projected error volume, the documented tendency of AI detectors to flag text by non-native English speakers, and the fact that Turnitin gives no detailed information as to how it determines whether a piece of writing is AI generated. Whether it is switched on where you study is a per-institution toggle that moves in both directions: NC State's DELTA unit published on 2 September 2025 that Turnitin AI Detection is re-enabled there. Ask your instructor whether the AI indicator is enabled on your specific assignment, which is a separate setting from whether your university licenses Turnitin at all. Checked 28 August 2026.
It is a per-institution toggle, and institutions move it in both directions. NC State's DELTA teaching unit published on 2 September 2025 that Turnitin AI Detection is re-enabled and available through the Turnitin Assignment tool, having previously been off. So the honest answer to whether your work is being scored for AI is that nobody on the internet can tell you; your institution's teaching and learning centre or your instructor can. Ask in that order. The concrete question is whether the AI writing indicator is enabled on the assignment you are about to submit to, not whether your university has a Turnitin contract, because those are two separate settings.
What false positive rate has Turnitin actually published, and has it changed?

April 2023, launch. Turnitin claims a false positive rate of less than 1 percent. This is the number Vanderbilt was reacting to.
23 May 2023. Chief product officer Annie Chechitelli reports 38.5 million submissions processed, 9.6 percent flagged over 20 percent AI writing and 3.5 percent at 80 to 100 percent, and announces three changes made in response to false positive complaints: an asterisk instead of a percentage for documents under 20 percent AI writing, a minimum word count raised from 150 to 300, and adjusted sentence aggregation at the start and end of a document.
14 June 2023. Turnitin publishes the sentence level figure, and it is four times worse than the headline. The document level rate of less than 1 percent applies only to documents already scoring 20 percent or more AI writing. The false positive rate for individual highlighted sentences is around 4 percent, and Turnitin adds that 54 percent of the time those false positive sentences sit directly next to actual AI writing.
24 February 2026. Turnitin reports that since October 2025 roughly 15 percent of essay submissions contained more than 80 percent AI generated writing, up from about 3 percent when the detector launched in April 2023.
Put plainly: the 1 percent figure is a document level rate, measured on documents already scoring 20 percent or higher, on Turnitin's own data. It is not a per sentence rate and it is not independently audited.
What do independent studies find when they test AI detectors?
Two peer reviewed studies are worth knowing by name, and both are less flattering than any vendor page. Weber-Wulff and colleagues tested 12 public tools plus Turnitin and PlagiarismCheck in the International Journal for Educational Integrity in 2023 and concluded that the available detection tools are neither accurate nor reliable and have a main bias towards classifying the output as human written. Perkins and colleagues ran 805 samples through six detectors for the International Journal of Educational Technology in Higher Education in 2024 and found mean accuracy of 39.5 percent on unmanipulated AI generated text, with only 67 percent of the human written control samples correctly identified as human. Their baseline ranking put Copyleaks first at 64.8 percent, Turnitin second at 61 percent and GPTZero last at 26.3 percent.
One caveat matters more than the numbers: both studies tested models that no longer exist. Turnitin added AI paraphrasing detection on 16 July 2024, which highlights text the model predicts was AI written and then run through a paraphrasing tool in a different colour from plain AI writing. It shipped AI bypasser detection on 27 August 2025, English only at launch. Turnitin Clarity, generally available since 15 July 2025, moved the company's own emphasis toward showing drafting history rather than a probability. All three post-date the data above, so treat any accuracy percentage you read, these included, as a dated observation rather than a current property of the product. Checked 28 August 2026.
Why do non-native English writers get flagged more often?
Because low perplexity, meaning predictable word choice, is what these models score, and careful second-language writing is predictable by construction. Liang and colleagues at Stanford tested seven public GPT detectors on 91 human written TOEFL essays and 88 US eighth grade essays. Accuracy on the eighth grade essays was near perfect. On the TOEFL essays the average false positive rate was 61.22 percent, all seven detectors unanimously flagged 18 of the 91, and 89 of the 91 were flagged by at least one. Asking ChatGPT to enhance the word choices to sound more like that of a native speaker cut that average to 11.77 percent on the same essays.
Turnitin was not one of the seven detectors tested, so this is evidence about perplexity based detectors as a class rather than a measurement of Turnitin. It is still the clearest published demonstration of why the same essay reads as human or machine depending only on how ornate its vocabulary is.
Does using Grammarly make Turnitin flag your work?
It depends which half of Grammarly you used, and Turnitin's guidance draws the line explicitly: the detector is not tuned to target Grammarly generated spelling, grammar, and punctuation modifications, but content produced using Grammarly's generative features will likely be flagged as AI generated. Fixing a comma splice is not the trigger. Asking a tool to write the sentence is. Check which features you have switched on before deciding whether a flag surprised you.
More questions about Turnitin's accuracy
Has anyone independently tested Turnitin's current 2026 model?
Not that we can point to. The peer reviewed tests worth citing were run in 2023 and 2024, before paraphrasing detection and bypasser detection shipped, and Turnitin does not sell individual access, which makes independent replication awkward by design. Anyone quoting a precise 2026 accuracy figure for Turnitin should be asked where they got a licence and when they ran it.
How this was made: this post predates this site's "How this was made" disclosure convention, which was added 2026-08-16. Drafting was AI-assisted with human editing. This paragraph and the diagram above were added during a 2026-08-29 accuracy pass; the diagram summarises points the post already made, and no claim in this post was re-verified beyond what is stated above.


