Why non-native English writing gets flagged as AI
A 2023 Stanford HAI study found that AI detectors misclassify non-native English writing as AI-generated at far higher rates than native writing. Here is what the study found, why the bias exists, and what non-native writers can do about it.The Stanford HAI finding
The Stanford HAI study, titled "GPT detectors are biased against non-native English writers," tested seven popular AI detectors on essays from both native and non-native English speakers. It found that the detectors classified non-native essays as AI at dramatically higher rates, in some samples flagging the majority of essays, even though they were human-written.The original announcement is available at Stanford HAI, and the full paper is published on arXiv at arxiv.org/abs/2304.02819. It has become the reference citation whenever detector bias is discussed.Why the bias exists (the mechanism)
AI detectors score statistical patterns. The specific patterns they look for, low perplexity, low burstiness, limited vocabulary variance, formulaic sentence structure, are features of AI writing. They are also features, coincidentally, of how many non-native English speakers write:- Vocabulary range tends to be narrower. Non-native writers often stick to vocabulary they are confident with, leading to statistically similar word choice patterns.
- Sentence structure tends to be more formulaic. Learned grammatical templates produce lower burstiness scores.
- Register tends toward formal or mid-formal. School-taught English skews toward the same register models default to.
Why it is hard to fix
A few fixes have been proposed, none cleanly working.Retrain on more diverse corpora. The natural answer, but it requires labeled training data at scale. Getting a large, diverse corpus of non-native English that is reliably human-authored, and matched controls from native speakers and AI, is expensive. Progress is being made but slowly.Calibrate differently for non-native writers. This requires knowing in advance whether a writer is native or non-native, which detectors usually do not. And it raises ethical questions: a detector should not need to know a writer's background to classify their text fairly.Reduce reliance on statistical signals. More specific signals help, but are not a complete answer, because distinctive patterns can still overlap with non-native writing styles in some cases.What current detectors are doing about it
Some detectors now publish bias audits. Some calibrate more conservatively, raising the score threshold for flagging, which reduces false positives but also raises false negatives. Some provide explanations, such as per-sentence signal breakdowns, that help writers and reviewers understand why a score came back the way it did.Leap's free AI score checker highlights the sentences that weigh most in the score, so you can see whether a high score comes from vocabulary-variance and rhythm signals rather than anything you actually did. It runs in your browser, nothing is sent anywhere, and it is a signal, not proof.What this means for non-native writers
Three practical things:- Keep drafts. If your paper gets flagged, your revision history, Google Docs, Word, git, is the strongest defense. A real draft trail proves human authorship in a way no detector can counter.
- Cross-check. Run the text through multiple detectors. If they disagree, the statistical signal is not consistent, which supports your position if challenged.
- Cite the Stanford finding. If you are in an academic context and your honest writing gets flagged, the Stanford HAI paper is the canonical reference. Including it in a response elevates the conversation from "did the detector catch you?" to "does this detector have a known bias?"