How AI detectors actually work
AI detectors do not read your text the way a person does. They score statistical patterns in the writing, and this guide explains what those patterns are and why they matter.The core signals
Every modern AI detector scores variations of the same core signals. The weights and implementations differ between products, but the underlying ideas are the same: how predictable the words are, how uniform the sentences are, and how the vocabulary is distributed.1. Perplexity (how predictable is the next word?)
A language model produces text by picking likely next tokens given the previous ones. When you read AI output, each word is roughly what the model thought should come next. Measured against a language model, that sequence scores as highly predictable: low surprise, low perplexity.Human writing tends to score higher perplexity. People pick unexpected words, odd phrasings, specific references, and rhetorical moves that models do not prioritize. Low perplexity is one of the strongest signals of AI authorship.2. Burstiness (do sentence lengths vary?)
Human writing is bursty. It alternates between short, punchy sentences and long, complex ones. AI writing tends to land in the middle and stay there: consistent medium-length sentences with similar syntactic structure.Detectors look at the variation in sentence length and complexity. Low variance is a red flag. It is why a string of well-constructed AI sentences reads as AI even when each sentence individually looks human.3. Token probability distribution and vocabulary
Beyond per-word predictability, detectors score the shape of the whole word distribution. AI text uses a narrower vocabulary band, reaching for words like 'leverage', 'navigate', 'multifaceted', and 'ensure' at rates far higher than humans do. Heavy use of these stock AI terms is a strong signal.4. Phrasing habits
AI models share recognizable phrasing habits: hedging words, formulaic transitions, conversational preambles, and punctuation patterns like heavy em-dash use. Detectors that scan for these habits can flag text that otherwise looks natural on the surface.These habits are signals, not proof. A human who writes in a formal, even style can trip the same flags, which is why no detector should be treated as the final word.Why paraphrasing doesn't fool detectors
A naive paraphrase swaps synonyms and reorders clauses. That changes the surface but leaves the underlying distribution mostly intact:- Perplexity stays low, because the words are still ones a model would rank highly.
- Burstiness doesn't change, since sentence lengths and structures remain uniform.
- Vocabulary stays in the same stock AI band.
What a detector can't tell you
Every detector has limits. Here is what they cannot reliably know:- Short text (under roughly 50 words) does not give enough statistical signal to classify reliably.
- Highly technical prose, dense and term-heavy by subject matter, can look AI-like even when a human expert wrote it.
- Non-native English sometimes scores as AI because the rhythm and vocabulary patterns differ from native-English training corpora.
- Heavily edited AI text, where a human genuinely rewrote large portions, can land in a gray zone that is hard to call either way.