Start free

How accurate are AI detectors, really?

Detector accuracy claims come from benchmarks, but real-world performance on everyday text is often much weaker. Here is what the gap between marketing numbers and daily use actually looks like, and how to read a score responsibly.

Try Leap's free AI tools

Leap's AI detector runs in your browser, free, with no signup and nothing sent anywhere. It scores your text from 0 to 100 and highlights the sentences that weigh most, so you can see the signal for yourself.

How the marketing numbers are produced

Detector accuracy claims come from benchmarks the detector company controls. They build a corpus of human-written and AI-written text, run their detector, and report the percentage it classified correctly. The issue: the corpus is theirs. It is typically balanced (50/50 human and AI), uses model output from models the detector was trained on, and avoids edge cases.Real-world performance on out-of-distribution text, such as non-native English, highly technical prose, heavily edited AI output, or brand-new models released that week, is materially worse. Exact numbers vary by study, and no single published figure should be treated as definitive.

What the independent research shows

A 2023 Stanford HAI study tested seven popular detectors against non-native English TOEFL essays and found high false-positive rates, with the worst offender flagging the large majority of confirmed-human essays as AI.Follow-up research in 2024 tested detectors against student writing at scale and found that even the better-performing tools still produced false positives on confirmed-human essays. On a large student cohort, even a small false-positive rate means dozens of writers wrongly flagged.On the other side, research has shown that basic paraphrasing of AI output can sharply reduce detector accuracy across most tools. OpenAI's own classifier was withdrawn over exactly this kind of reliability problem. The honest conclusion: accuracy varies widely by text, model, and editing, and any single score should be treated cautiously.

Why detectors disagree with each other

Run the same text through three detectors and you will often get three different answers. Each detector:
  • Trains on a different corpus, with different weights for different writing signals
  • Uses different classifier thresholds, so one's 70% AI is another's 85%
  • Updates at different cadences, so newer-model support varies by weeks to months
  • Handles edge cases differently, such as non-native English, technical prose, or very short text
The disagreement is a feature of the space, not a bug in any one tool. Our piece on how AI detectors actually work explains why two classifiers trained on different data can look at the same sentence and produce different probabilities.

How to interpret a detector score responsibly

A single detector's score is a signal, not a verdict. Good practice:
  • Cross-check with more than one detector. If all agree at high confidence, the signal is stronger. Disagreement means gray zone.
  • Consider context. Non-native English writer? Technical subject? Short text? All three increase false-positive risk.
  • Look at per-sentence scores, not just the overall number. If one paragraph scores high and the rest score low, the issue is localized.
  • Never use a single score as evidence. In education, HR, or any consequential setting, detector output should be one input among many, never the only basis for accusing anyone.

What makes Leap's detection different

Leap runs entirely in your browser: free, no account, and nothing is sent anywhere. It returns:
  • An overall score from 0 to 100
  • Per-sentence highlighting that shows which sentences weigh most on the score
  • The writing signals behind the result: uneven versus uniform sentence length (burstiness), stock AI phrases, hedging and transition words, em-dash density, and repetition
The expanded output is what makes a score actionable. A single number says something looks off somewhere. Sentence highlighting plus the signal breakdown says these sentences trip the score because of this pattern, which is what you need to revise instead of guessing.Leap's score is a signal, not proof. It can be wrong, and it should never be the only basis for a decision about a person. For a deeper look at the edge cases that drive false positives, such as non-native English, technical prose, and short inputs, see our piece on why human writing gets flagged as AI.

Frequently asked questions