Start free

A brief history of AI writing detection

From academic stylometry to ChatGPT, GPTZero, and the modern detector market: the key dates and tools that shaped AI writing detection, and what the history tells us about where it is going.

Pre-2022: academic research

Distinguishing machine-generated text from human writing has been an academic research problem for decades, predating LLMs entirely. Classic work on statistical stylometry, identifying authorship through sentence-length variance, vocabulary distribution, and syntactic patterns, laid the foundation. arXiv has a long trail of pre-2022 papers on authorship attribution that read surprisingly relevant today.But as a consumer product category, AI detection did not exist. GPT-2 was released in 2019 without widespread concern; GPT-3 in 2020 raised more eyebrows; neither triggered a detection industry.

November 2022: ChatGPT launches

ChatGPT launched on November 30, 2022, and within two months had over 100 million users. Teachers, editors, and employers immediately needed a way to detect AI-written content. The detection industry was born in response.

January 2023: GPTZero

Edward Tian, then a Princeton senior, released GPTZero in early January 2023. It went viral almost immediately. Tian's approach combined perplexity (low perplexity equals AI) and burstiness (variance of sentence length) as the two primary signals. The core architecture set the template that every subsequent detector would use.

January 2023: OpenAI's AI Text Classifier

OpenAI released its own detector three weeks after GPTZero. It promised to identify AI-written text with reasonable accuracy, but came with disclaimers about false positives. The tool was hosted at platform.openai.com and free.

Spring 2023: the gold rush

Within six months, the market had Originality.ai, Copyleaks, Turnitin's AI indicator, ZeroGPT, Sapling, Writer's detector, and dozens of smaller tools. Each claimed higher accuracy than the others. Benchmarks were rare and inconsistent. Schools and universities started adopting them rapidly, and a simultaneous backlash emerged: false positives, bias against non-native English writers, and the fundamental limits of statistical detection.

July 2023: OpenAI withdraws its classifier

On July 20, 2023, OpenAI took down the AI Text Classifier, citing "low rate of accuracy." The public retraction was a concession that even the company with the most complete picture of how its own models wrote couldn't reliably detect their output. This was a meaningful moment: it told the market that perfect detection wasn't around the corner.

Late 2023: the bias studies

Researchers at Stanford HAI and elsewhere published papers showing that AI detectors misclassified non-native English writing as AI at alarming rates. The Stanford study became the most-cited evidence of detector bias. It forced the industry to reckon with limits the marketing had been ignoring.

2024: humanizers emerge

Undetectable.ai, QuillBot's humanizer, Phrasly, StealthGPT, and others launched products specifically designed to defeat AI detectors. The ensuing cat-and-mouse dynamic shaped the year. Detectors retrained; humanizers updated; detectors retrained again. Model fingerprinting emerged as a claimed signal, though how well it survives heavy paraphrase remains an open question.

2024: Turnitin integrates detection

Turnitin's AI writing indicator became bundled into its Originality product at scale, reaching most K-12 and university LMS deployments. This turned AI detection from a research curiosity into a routine workflow for teachers.

2025: watermarking and policy

OpenAI and DeepMind both reported work on watermarking, embedding invisible fingerprints in model output that detectors could reliably pick up. Implementation remained voluntary and inconsistent; watermarks can often be stripped by light editing.Policy frameworks matured in parallel. The EU AI Act came into force with provisions around AI-generated content labeling. Several US states introduced similar legislation. Detection became less a technical question and more a compliance question.

Where we are now

Current detectors layer multiple signals, including perplexity, burstiness, hedging density, em-dash patterns, and vocabulary distribution, to produce scores with explanations. The best tools are honest about false positives and provide signal breakdowns rather than single confidence numbers. Leap's free detector follows this approach: it runs in your browser, highlights the sentences that weigh most, and treats every score as a signal, not proof.The market has consolidated around GPTZero, Turnitin, Originality, Copyleaks, and a handful of newer entrants. The humanizer market has matched it, with Undetectable, QuillBot, StealthGPT, and others. Our piece on free detectors compares the current options.

What the history tells us

Three patterns stand out. First, detection never stays solved: every new model release requires retraining. Second, no detector is perfect; bias and false positives are structural, not bugs to be fixed. Third, the field is moving from pure technical detection toward a hybrid of disclosure, watermarking, and policy. If you're reading this in 2028, the specifics will have moved; the dynamics probably won't have.If you want to see what current detection looks like in practice, try Leap's free detector with your own text.

Frequently asked questions