
Merlin's AI detector is one feature inside a much larger AI assistant, not a standalone product built around detection. That distinction matters more than any accuracy number, because it changes what the tool was actually optimized for. Here is what the detector measures, how much you should trust its score, and how it stacks up against a detector your institution might actually use.
A bundled feature, not a specialist tool
Merlin's core product is a general AI assistant: writing, summarizing, browsing, chat. Detection is one capability among many, which means it competes for engineering attention against every other feature in the roadmap. That is not automatically a problem, but it is a real trade-off worth naming before you rely on the score for anything that matters. A tool whose entire business depends on detection accuracy has a different incentive to keep retuning against the latest models than a feature bolted onto a broader suite.
What the score is actually measuring
Like every detector in this category, Merlin is not reading a hidden watermark or a confession baked into the text. It is scoring statistical properties: perplexity, meaning how predictable each next word is, and burstiness, meaning how much sentence length and rhythm vary across a passage. Human writing tends to wander unevenly. Model output tends to stay smooth on both measures, because smoothness is what the underlying training objective rewards.
The mechanics are shared across every detector on the market, which is why understanding them once transfers everywhere. We break them down fully in how ChatGPT detectors work.
How accurate is it, honestly
No outside reviewer can hand you a trustworthy hard accuracy number, and any review that quotes one confidently is guessing. What is knowable is the shape of the errors, and that shape repeats across this entire category of tool.
Short passages are unreliable. Under roughly 300 words there isn't enough signal, and the score swings hard on small edits.
Non-native English writing is a known false-positive risk. Simpler vocabulary and more regular structure can look statistically machine-like even when a person wrote every word.
Technical and formulaic writing flags more often. Reports, summaries, and documentation are supposed to be uniform, and uniformity is exactly the signal detectors key on.
If Turnitin is the detector you're actually being graded against, the realistic accuracy picture for that specific tool is in how accurate is Turnitin AI detection, and the caveats above apply just as hard there.
What to do with a Merlin score
Treat it as one data point, not a verdict. If you're checking your own writing before submitting it somewhere else, run the same paragraph through a second, dedicated detector and see whether the scores agree. Agreement across tools built on different training data is a stronger signal than any single number.
Check your text with our free AI detector and compare it against Merlin's score before trusting either one alone.
If you need to lower a score, do it properly
Chasing a lower number with synonym swaps rarely holds up, because swaps don't touch the rhythm the detector is actually measuring. Real structural rewriting, varied sentence length, and a detail only you would know move the needle. Guessing at a target percentage does not.
The full mechanism behind why shallow edits fail is in why AI humanizers don't work (sometimes).
Our free AI Humanizer handles the structural pass if you need one.
Frequently asked questions
Is the Merlin AI detector accurate?
It measures the same statistical signals every detector in this category measures, with the same known blind spots on short, technical, or non-native-English text. Treat any single score as one data point, not a verdict.
Is Merlin's AI detector free?
Pricing and free-tier limits for bundled AI assistants change often, so check current terms directly before relying on it for repeated checks.
Should I use Merlin instead of a dedicated AI detector?
For a quick gut check, it's fine. For anything graded or high-stakes, cross-check with a second, dedicated detector rather than trusting one bundled feature's number alone.


