Are AI Detectors Accurate? What the Research and Our Tests Show
AI detectors catch raw ChatGPT text but miss edited AI and flag some human writing. Here is what studies and our own tests show, and how to use them fairly.

AI detectors are accurate at spotting plain, unedited ChatGPT output, but they can't prove beyond doubt that a particular person used AI. Independent studies have shown that widely used detectors miss much of the AI text that has been edited, and wrongly flag a lot of human writing, especially writing by non-native English speakers. OpenAI even discontinued its own AI text classifier because it wasn't accurate enough. Treat a detector score as a reason to look closer, never as proof.
Key takeaways
- Detectors are reliable on long, unedited AI text and unreliable on short, edited, mixed or formulaic text.
- OpenAI's own classifier caught only 26% of AI-written text and was withdrawn in July 2023.
- A Stanford study found detectors flagged 61% of essays by non-native English writers as AI-generated.
- Different detectors often disagree on the same text. In our tests, one rewritten passage scored 10%, 25% and 78% AI on three tools.
- Use a detector to decide where to look, then check drafts, sources and version history before drawing any conclusion.
The short answer: when AI detectors work and when they don't
Most AI detectors look for text that is statistically predictable and unnaturally even in rhythm, which is how a language model writes by default. That makes them good at one thing and much less reliable at most others.
| Situation | How reliable detectors are |
|---|---|
| Long, unedited ChatGPT, Claude or Gemini output | High. Usually flagged at 90% to 100% AI |
| AI text that a person edited heavily | Low to medium. Scores vary a lot between tools |
| AI text run through a paraphraser | Medium. Simple synonym swaps are often still caught |
| AI text rewritten by a strong humanizer | Low. Some tools score it as human |
| Mixed human and AI writing | Low. Hard to tell where one ends and the other begins |
| Short text under about 250 words | Low. Too little text to measure patterns |
| Formal, technical or formulaic human writing | Risky. A common source of false positives |
| Writing by non-native English speakers | Risky. Simpler phrasing can look machine-made |
So the honest answer to "are AI detectors accurate?" is: they're reliable against low-effort AI writing and fall short when the stakes are highest, such as a student essay that was partly edited, or a professional who simply has a plain, consistent writing style.
What the research says about AI detector accuracy
Detector companies often advertise accuracy of 98% or 99%. Those numbers usually come from their own tests on clean samples of fully human and fully AI text. Independent research paints a more mixed picture.
OpenAI withdrew its own detector
In January 2023, OpenAI released a free AI text classifier. By its own figures, the tool correctly identified only 26% of AI-written text as "likely AI-written" and wrongly labeled human text as AI 9% of the time. OpenAI withdrew it in July 2023, citing its low rate of accuracy. If the company that built ChatGPT couldn't reliably detect ChatGPT, that says a lot about how hard the problem is.
Detectors are biased against non-native English writers
A 2023 Stanford study by Liang and colleagues, published in the journal Patterns, ran essays through seven popular AI detectors. The detectors classified essays written by US eighth-graders almost perfectly. But they flagged 61% of TOEFL essays written by non-native English speakers as AI-generated on average, and nearly all of those essays were flagged by at least one detector.
The reason is that non-native writers tend to use a smaller vocabulary and more common sentence patterns. That makes their writing statistically predictable, which is exactly what detectors are trained to treat as a sign of AI. We explain those signals in how AI detectors work: perplexity and burstiness.
A 14-tool test found none were reliable
In 2023, a European team led by Debora Weber-Wulff looked at 14 AI detection tools, including Turnitin and several free ones, and reported in the International Journal for Educational Integrity that they were neither accurate nor reliable. None of the tools reached 80% accuracy, and they tended to label text as human. Performance dropped further when AI output was paraphrased or machine-translated.
Turnitin's own numbers
Turnitin, the detector most schools use, says its document-level false positive rate is below 1% for papers where more than 20% of the text is flagged as AI. At the sentence level, it has reported a false positive rate of around 4%. That's why Turnitin hides scores under 20% and tells instructors not to rely on the score alone. Read more in our guide to the Turnitin AI checker.
An error rate of 1% doesn't seem like much until you scale it up. When Vanderbilt University switched off Turnitin's AI detector in August 2023, it noted that 1% of the roughly 75,000 papers it had submitted the previous year would have meant hundreds of students falsely flagged. Several other universities did the same.
Why AI detectors get it wrong
Detectors don't know who wrote a text. They only calculate the probability that a language model would have produced it. Several things can skew that calculation:
- Predictable human writing. Formal reports, legal text, technical documentation and famous historical texts all use predictable phrasing. Some detectors have even labeled the US Constitution as AI-written, because it appears so often in training data that it looks "predictable" to a model.
- Grammar and rewriting tools. Text polished by Grammarly or a paraphraser can pick up the smooth, even rhythm that detectors associate with AI.
- Short samples. A 100-word paragraph doesn't give a detector enough sentences to measure variation. Most tools warn that results under a few hundred words are unreliable.
- Editing and humanizing. Once a person or a humanizer reworks AI text, its statistical fingerprint changes. Detectors then have to guess.
- New models. Detectors are trained on output from existing models. Each new model writes a little differently, and detectors take time to catch up.
Our tests: detectors disagree on the same text
The HumanizerPRO team has run dozens of detector tests while reviewing AI tools. Two results stand out.
First, false positives happen with real, human-written content. In July 2025, Originality.ai rated a human-written blog post from 2022, published before ChatGPT existed, as 60% likely AI. That was on its Lite model, which is designed to reduce false positives. Full details are in our Originality.ai review.
Second, detectors disagree with one another. When we checked a passage rewritten by StealthWriter, GPTZero scored it about 10% AI, StealthWriter's own checker 25% and Originality.ai 78%, all for the very same text. See the full test in our StealthWriter review.
We've also seen the opposite: not even our own humanizer beats every detector every time. HumanizerPRO's Standard mode output scored 71% AI on Originality.ai in one test, while a text humanized with HumanizerPRO scored 0% AI on Turnitin in another. Detectors are built differently, tuned differently and updated on different schedules, so a result from one detector doesn't tell you what another will say. We look at how the two sides are evolving in AI humanizers vs. AI detectors.
How to use an AI detector responsibly
A detector is useful when it's one input among several, alongside a close read for the signs of AI writing. Whether you're a teacher, an editor or checking your own writing, these habits make the results far more meaningful:
- Check at least 300 words. Longer samples give more reliable scores.
- Look at the highlighted sentences, not just the percentage. A good detector shows which lines read as AI, which tells you far more than one number.
- Use more than one detector. If tools strongly disagree, the text is in a gray zone and no single score should be trusted.
- Consider the writer. Non-native English speakers and writers of formal or technical text are more likely to be flagged unfairly.
- Ask for the process, not just the product. Drafts, notes, outlines and Google Docs or Word version history show how a piece was written.
- Check for other problems too. AI misuse often shows up as invented sources or copied passages, so run a plagiarism check and a fact check as well.
- Never act on a score alone. Turnitin, GPTZero and most other vendors say the same thing in their own guidance.
See which sentences read as AI. HumanizerPRO's free AI detector scores text sentence by sentence, so you can see exactly what a detector is reacting to instead of guessing from a single number.
Which AI detector is the most accurate?
There's no single best detector, because each one strikes its own balance between catching AI and avoiding false positives.
| Detector | Strictness | Best for | Free option |
|---|---|---|---|
| Turnitin | Conservative, hides scores under 20% | Schools and universities | No, institutions only |
| GPTZero | Moderate | Teachers and students | Yes |
| Originality.ai | Very strict, more false positives | Publishers and agencies | No, paid credits |
| Copyleaks | Strict | Businesses and schools | Limited |
| HumanizerPRO AI detector | Moderate, sentence-level highlights | Checking your own writing before you submit or publish | Yes |
If you must catch as much AI as you can and can tolerate a few false positives, a strict tool like Originality.ai will work. If a wrongful accusation could harm someone, a cautious tool plus your own judgment is the more prudent choice. For checking your own work, a free detector with sentence highlights shows you what to revise.
What to do if a detector flags your writing
If your own work is flagged, don't panic. A detector score is not proof. Ask to see which passages were highlighted, gather your drafts and version history, and explain how you researched and wrote the piece. If English isn't your first language, or the writing is formal by nature, mention that detectors are known to struggle with that kind of text. Our guide to what to do if you're falsely accused of using AI walks through each step, with an email template you can adapt.
If you used AI in any part of the draft and your school or client permits it, disclose it. And if your natural writing keeps getting flagged, more varied sentence length and more concrete detail usually fix most of the problem. Some writers also use an AI humanizer to refine AI-assisted drafts, as we explain in what AI humanization is and when it helps.
Frequently asked questions
Are AI detectors accurate?
To a certain extent. They're good at detecting long, unedited AI text, but they miss a lot of edited or humanized AI writing and can wrongly flag perfectly normal human writing. Independent studies have found real-world accuracy much lower than the 98% to 99% many vendors advertise.
Can AI detectors be wrong about human writing?
Yes. There are many documented cases of detectors wrongly flagging human work. A Stanford study found detectors labeled 61% of essays written by non-native English speakers as AI, and in the HumanizerPRO team's July 2025 test, Originality.ai rated a human-written 2022 blog post as 60% AI.
Which AI detector is the most accurate?
No detector is best in every case. Strict tools like Originality.ai catch more AI text but produce more false positives, while Turnitin is more conservative. Using two or more detectors and reviewing the highlighted sentences gives a more reliable picture than any single score.
Why was my own writing flagged as AI?
Usually because the wording and sentence length are easy to predict. Formal or technical writing, grammar-checked text, short samples and writing by non-native English speakers are the most common sources of false positives.
Can AI detectors detect ChatGPT?
Yes, unedited ChatGPT output is usually detected with a high score. Once the text has been edited by a person, humanized, or mixed with human writing, it is much less likely to be picked up.
Do AI detectors work on short text?
Not well. Most detectors need at least a few hundred words to measure patterns reliably, and many warn that results on short passages are unreliable.
Is there a free AI detector?
Yes. HumanizerPRO (humanizerpro.ai) is a free AI humanizer with a built-in AI detector, plagiarism checker and fact checker. Its detector highlights AI-sounding text sentence by sentence. GPTZero also offers a free tier.

Kamran Khan
Kamran Khan is the founder of HumanizerPRO and a leading voice in the ethical use of AI-generated content. With years of hands-on experience in AI, SEO, and digital publishing, he built HumanizerPRO to help creators and professionals turn robotic AI text into clear, human-like writing that meets real-world standards.