How accurate are AI detectors, really?
Detector accuracy claims come from benchmarks, but real-world performance on everyday text is often much weaker. Here is what the gap between marketing numbers and daily use actually looks like, and how to read a score responsibly.Try Leap's free AI tools
Leap's AI detector runs in your browser, free, with no signup and nothing sent anywhere. It scores your text from 0 to 100 and highlights the sentences that weigh most, so you can see the signal for yourself.How the marketing numbers are produced
Detector accuracy claims come from benchmarks the detector company controls. They build a corpus of human-written and AI-written text, run their detector, and report the percentage it classified correctly. The issue: the corpus is theirs. It is typically balanced (50/50 human and AI), uses model output from models the detector was trained on, and avoids edge cases.Real-world performance on out-of-distribution text, such as non-native English, highly technical prose, heavily edited AI output, or brand-new models released that week, is materially worse. Exact numbers vary by study, and no single published figure should be treated as definitive.What the independent research shows
A 2023 Stanford HAI study tested seven popular detectors against non-native English TOEFL essays and found high false-positive rates, with the worst offender flagging the large majority of confirmed-human essays as AI.Follow-up research in 2024 tested detectors against student writing at scale and found that even the better-performing tools still produced false positives on confirmed-human essays. On a large student cohort, even a small false-positive rate means dozens of writers wrongly flagged.On the other side, research has shown that basic paraphrasing of AI output can sharply reduce detector accuracy across most tools. OpenAI's own classifier was withdrawn over exactly this kind of reliability problem. The honest conclusion: accuracy varies widely by text, model, and editing, and any single score should be treated cautiously.Why detectors disagree with each other
Run the same text through three detectors and you will often get three different answers. Each detector:- Trains on a different corpus, with different weights for different writing signals
- Uses different classifier thresholds, so one's 70% AI is another's 85%
- Updates at different cadences, so newer-model support varies by weeks to months
- Handles edge cases differently, such as non-native English, technical prose, or very short text
How to interpret a detector score responsibly
A single detector's score is a signal, not a verdict. Good practice:- Cross-check with more than one detector. If all agree at high confidence, the signal is stronger. Disagreement means gray zone.
- Consider context. Non-native English writer? Technical subject? Short text? All three increase false-positive risk.
- Look at per-sentence scores, not just the overall number. If one paragraph scores high and the rest score low, the issue is localized.
- Never use a single score as evidence. In education, HR, or any consequential setting, detector output should be one input among many, never the only basis for accusing anyone.
What makes Leap's detection different
Leap runs entirely in your browser: free, no account, and nothing is sent anywhere. It returns:- An overall score from 0 to 100
- Per-sentence highlighting that shows which sentences weigh most on the score
- The writing signals behind the result: uneven versus uniform sentence length (burstiness), stock AI phrases, hedging and transition words, em-dash density, and repetition