Perplexity vs burstiness: why AI detectors flag text
Two signals sit behind most AI detection: how predictable your word choices are, and how much your sentence lengths vary. Here is what each one measures and why both matter.Try Leap's free AI tools
Leap's detector runs entirely in your browser: free, no signup, and your text is never sent anywhere. Paste a draft and see the score in under a minute. Try Leap on your own draft.Perplexity
Perplexity measures how surprised a language model is by a given text. A model reads your sentence, tries to predict each word based on the prior words, and scores how 'unexpected' each actual word was. High perplexity means the model found the text surprising. Low perplexity means the model thought the text was predictable.AI-generated text scores low perplexity by construction. The model wrote it; the words are exactly what the model expected. When a detector runs the text back through a language model, the probability scores come out unusually smooth, with each word fitting the expected distribution.Human writing has more spikes. Humans pick words for reasons a model does not optimize for: memory, rhythm, specificity, humor, context only the writer has. Those choices show up as perplexity spikes and flag as human-authorship signals.Burstiness
Burstiness measures variation in sentence structure over a passage. The classic metric: standard deviation of sentence length, plus standard deviation of syntactic complexity. High burstiness means aggressive alternation between short punchy sentences and long complex ones. Low burstiness means consistent sentence length and structure.Human writing, especially effective human writing, is bursty. A three-word sentence. Then a long, winding sentence that picks up momentum and drives to a conclusion with multiple dependent clauses. Then another short one. This rhythm is what makes writing feel natural.AI writing tends to be smoother. Medium length, consistent structure, parallel clauses. Even when AI produces a short sentence, the surrounding sentences are similar in length. Low burstiness is a strong AI signal.Why you need to address both
Naive editing strategies fix one signal and leave the other intact:- Synonym-swapping paraphrase changes individual words (barely affects perplexity, because synonyms are also high-probability) and does not touch burstiness at all.
- Sentence reordering changes neither perplexity nor burstiness; you are just shuffling the same low-perplexity, low-burstiness distribution.
- Adding filler sometimes raises burstiness (filler sentences are often short) but lowers perplexity further, because filler is maximally predictable.