UndetectableTest

Updated: 2026-09-28

Guide

How AI Detectors Work: Perplexity, Burstiness & Limits (2026)

AI detectors don't read your writing the way a teacher does. They run statistics on it — measuring how predictable your word choices are and how uniform your sentences sound. Understand those two measurements and you understand every detector on the market, including why they all fail in the same ways.

The core idea: statistics, not understanding

A language model generates text by repeatedly picking likely next words. That leaves a statistical fingerprint: the output hugs the probable, avoids the surprising, and keeps a steady rhythm. Detectors invert the process — they score how “model-like” a text's statistics are. No comprehension is involved, which is both why detectors work at all and why they're so easy to fool: anything that disturbs the statistics disturbs the verdict.

Perplexity, explained simply

Perplexity measures how predictable word choices are. When a model writes, it favors the most likely next word — “the cat sat on the mat” rather than “the cat sat on the radish.” Sustained low perplexity (everything predictable) reads as AI; spikes of unpredictability read as human. Human writers surprise constantly — unusual metaphors, odd word choices, domain jargon used loosely. That's the signal.

The catch: plenty of human writing is low-perplexity too. Legal boilerplate, technical documentation, ESL writing with simple vocabulary, and the five-paragraph school essay all hug predictable patterns — which is why detectors disproportionately flag non-native writers and formulaic genres. Low perplexity means “predictable,” not “artificial,” and detectors can't tell the difference.

Burstiness, explained simply

Burstiness measures variation in sentence structure — length, complexity, rhythm. Human writing is bursty: a long winding sentence, then a fragment. Then another medium one. AI output tends toward metronomic uniformity: every sentence a comfortable medium length with similar construction. Detectors quantify that variance, and low variance scores as AI.

This is why the most effective humanizing technique is structural, not lexical. Swapping synonyms doesn't change sentence rhythm; breaking sentences up, merging them, and varying cadence does. It's also why genuinely rewriting in your own voice beats every tool — your natural rhythm is already bursty.

Other signals detectors use

Why detectors fail: the five classic cases

Can detectors be trusted?

As screening signals, yes — with the right tool, the right content type, and sensible thresholds, they catch plenty of raw AI output. As proof, no — and no serious vendor claims otherwise. The responsible pattern: high thresholds, agreement across two tools, and human judgment before any accusation. Schools that skip those steps manufacture false accusations; publishers that skip them burn freelancer relationships.

What this means for you

What is perplexity in simple terms?

How predictable the word choices are. AI picks likely words, so its text has low perplexity; humans surprise more. Detectors flag sustained low perplexity — which also catches predictable human writing.

What is burstiness?

How much sentence length and structure vary. Humans mix long sentences, short ones, and fragments; AI tends toward uniform medium sentences. Low variation scores as AI.

Can detectors tell which AI model wrote something?

Some claim model attribution, but reliability is low — especially after editing. Treat attribution claims skeptically.

Do detectors work on short text like tweets?

Poorly. Under ~150 words there isn't enough statistical signal, and scores become unreliable. Don't trust short-text verdicts.

Why was my human writing flagged as AI?

Most likely: it's short, formulaic, or uses simple predictable phrasing — all of which mimic AI statistics. It's a false positive, and it's common. Keep your drafts as evidence and dispute it.

Will detectors get better?

They retrain constantly, and each generation closes some gaps — while humanizers open new ones. Expect the arms race to continue, not a final winner.

A brief history of AI detection

Detection started as a curiosity in the GPT-2 era (2019), when OpenAI itself released a detector for its own model — it worked because one lab controlled the generator. The ChatGPT explosion of late 2022 changed everything: suddenly everyone could generate fluent text, dozens of detectors launched within months, and accuracy claims got wild. 2023–2024 brought the education panic — schools banning, then un-banning, then regulating AI — and detectors became disciplinary infrastructure before the science was ready. 2025 was the counter-offensive: Turnitin's model update specifically targeted humanizer evasion, humanizer vendors answered with structural rewriting, and the “99% accuracy” marketing era quietly ended as false-positive scandals piled up. 2026: the mature arms race — no settled winner, process-based verification rising, and policy finally catching up to technology.

The lesson of that history: every confident claim about detection has eventually been humbled by the next model generation — on both sides. Treat current capabilities as temporary.

Glossary: detection terms in plain English

What's the difference between AI detection and plagiarism detection?

Plagiarism detection matches text against existing sources — it's lookup. AI detection guesses at statistical origin — it's inference. A text can be 100% original and still flag as AI, or plagiarized yet score 'human.' They're different tools for different questions.

Are detectors biased against non-native speakers?

Effectively, yes — not by intent but by statistics. Simpler, more predictable phrasing (common in ESL writing) mimics the low-perplexity patterns detectors flag. Independent tests consistently show elevated false-positive rates. It's the industry's most serious fairness problem.

What detectors can't do (and won't soon)

Internalize that list and you'll read detector marketing — and detector scores — with appropriately calibrated skepticism. The technology is useful the way a smoke alarm is useful: good at alerting, terrible at convicting.

Can AI detect its own writing reliably?

Ironically, models are mediocre at recognizing their own outputs — detection models are trained separately on human-vs-AI classification, and self-recognition isn't a designed capability. Don't ask a chatbot 'did you write this' and trust the answer.

Will quantum computing change detection?

No — detection is a statistics and data problem, not a compute problem. The arms race will be decided by training data and model design, not raw hardware.

What's the single best defense against false accusation?

Process evidence: drafts, version history, notes. Statistical arguments about scores are weak; a document history showing the work being written is strong.