Start free

Perplexity vs burstiness: why AI detectors flag text

Two signals sit behind most AI detection: how predictable your word choices are, and how much your sentence lengths vary. Here is what each one measures and why both matter.

Try Leap's free AI tools

Leap's detector runs entirely in your browser: free, no signup, and your text is never sent anywhere. Paste a draft and see the score in under a minute. Try Leap on your own draft.

Perplexity

Perplexity measures how surprised a language model is by a given text. A model reads your sentence, tries to predict each word based on the prior words, and scores how 'unexpected' each actual word was. High perplexity means the model found the text surprising. Low perplexity means the model thought the text was predictable.AI-generated text scores low perplexity by construction. The model wrote it; the words are exactly what the model expected. When a detector runs the text back through a language model, the probability scores come out unusually smooth, with each word fitting the expected distribution.Human writing has more spikes. Humans pick words for reasons a model does not optimize for: memory, rhythm, specificity, humor, context only the writer has. Those choices show up as perplexity spikes and flag as human-authorship signals.

Burstiness

Burstiness measures variation in sentence structure over a passage. The classic metric: standard deviation of sentence length, plus standard deviation of syntactic complexity. High burstiness means aggressive alternation between short punchy sentences and long complex ones. Low burstiness means consistent sentence length and structure.Human writing, especially effective human writing, is bursty. A three-word sentence. Then a long, winding sentence that picks up momentum and drives to a conclusion with multiple dependent clauses. Then another short one. This rhythm is what makes writing feel natural.AI writing tends to be smoother. Medium length, consistent structure, parallel clauses. Even when AI produces a short sentence, the surrounding sentences are similar in length. Low burstiness is a strong AI signal.

Why you need to address both

Naive editing strategies fix one signal and leave the other intact:
  • Synonym-swapping paraphrase changes individual words (barely affects perplexity, because synonyms are also high-probability) and does not touch burstiness at all.
  • Sentence reordering changes neither perplexity nor burstiness; you are just shuffling the same low-perplexity, low-burstiness distribution.
  • Adding filler sometimes raises burstiness (filler sentences are often short) but lowers perplexity further, because filler is maximally predictable.
Real revision targets both. You raise perplexity by making specific word choices a model would not default to: unusual verbs, specific nouns, casual phrasings, intentional oddities. You raise burstiness by aggressively varying sentence length and breaking parallel structures.

How Leap surfaces these signals

Leap's free detector scores your text 0-100 from writing signals that include burstiness (uneven versus uniform sentence length), stock AI phrases, hedging and transition words, em-dash density, and repetition. It highlights the sentences that weigh most on the score, so you can see exactly which lines look the most uniform or predictable.That per-sentence highlighting is the actionable feedback missing from a bare score. If you want to understand what detectors generally do with signals like these, see our deeper explainer on how AI detectors actually work. The short version: humans vary; models do not.

The takeaway

You cannot address perplexity-and-burstiness detection with synonym swapping. You have to rewrite at the distribution level. That means specific word choices that models do not default to, sentence-length variation that exceeds AI baselines, and structural variety that breaks the parallel-clause habit most LLMs have.Leap's humanizer helps with the mechanics of that rewrite: it replaces stock AI phrases, removes em dashes and invisible characters, adds contractions, and marks sentences of uniform length so you can vary them yourself. It does not guarantee any detector result; the judgment and rewriting stay with you. For the practical workflow, see our guide on how to humanize AI text.

Why detectors still miss things

Perplexity and burstiness are strong signals, not perfect ones. Technical writing, legal boilerplate, and API documentation score naturally low on burstiness because the genre rewards uniform sentence structure. Non-native English prose can also cluster low on burstiness, which is one reason these writers get flagged more often.No score is proof of anything. Detection results can be wrong in both directions, and a score should never be the only basis for accusing anyone of using AI. For the full picture of who gets miscategorized, read our false positives deep dive, the flip side of this page.

Frequently asked questions