Every company selling AI detection to educational institutions leads with a number. Turnitin, GPTZero, Copyleaks, and Pangram all publish one, usually an accuracy rate or a false-positive rate, often reported to two decimal places. Pangram, for example, self-reports a false-positive rate of 0.19 percent and a false-negative rate of 1.4 percent on standard datasets. A false-positive rate of 0.19 percent means that out of 1,000 human-written essays, the detector falsely flags only about two texts as AI-generated.
Sounds great, doesn’t it? But unfortunately, none of those numbers can really tell you what will happen in your classroom. Because outside of a controlled lab environment, they have very little meaning.
I have made this point on The Augmented Educator Substack for over two years now, particularly in The Paradox of AI Detection. But so far I have mostly asserted rather than explained what might be the most complex part of this argument: that the trade-off at the center of these tools is subject to a hard mathematical limit. Better engineering can shift that trade-off. But nothing can eliminate it.
Absolutely nothing. The math is baked in.
So in this post, I want to complete my core argument against AI detection. And I have decided to do it in plain language and with no complex math notation. I wanted to keep this approachable to a lay audience. If you are interested in the finer details of this argument, you might want to check out the audio deep dive podcast episode linked at the end of this post.
How a score becomes a verdict
Let’s start with what a detector actually returns in practice. Underneath the percentages and the categories, it is fundamentally a score. The detector assigns a text passage a number between zero and one, say 0.83 or 0.11, showing how strongly the system associates it with machine-written text.
Early detectors often used that score to return a simple yes/no verdict, which ended up being highly problematic because it outsourced the integrity decision from the instructor to the detector. Most of the products available today therefore stop short of that. They present a percentage or a category and leave the final judgment of what that means to whoever reads the report. This separates the detector’s result from any misconduct accusation. It is now the teacher who makes that call and not the AI detector.
Turnitin, for example, reports what share of a document it estimates was AI-written. It simultaneously cautions that the figure cannot be used as proof of misconduct. GPTZero splits a submission into percentages of AI, mixed, and human. And Pangram sorts text into human, AI-edited, and fully AI-generated, and its EditLens model estimates how much AI editing a passage received.
None of that is an accusation. Every one of those reports leaves somebody else to decide what happens next, and that decision is a cutoff.
Somewhere between an innocent detection report reading some percentage and an email accusing a student of misconduct, a line gets drawn. Statisticians call that line a threshold. Think of it as a dial which the teacher sets, and think of the number it points at as a bar the score has to clear before anyone acts on it. Turn the dial up, and fewer texts clear the bar. The same detector, on the same essay, leads to entirely different outcomes depending on what number that dial points at.
When a company says its tool is 99.85 percent accurate, it is describing the result of applying a chosen threshold to a chosen test set. Move the dial or change the test set, and the advertised accuracy and error rates move with it.
Two piles that overlap
Let’s assume we want to use a detector to score an extensive set of known examples. Take thousands of texts you know a human wrote, for example, and score them all. Then take thousands you know came from a language model and score those too. You get two piles of scores.
If those piles sat in separate ranges, say, with every human text below 0.4 and every machine text above 0.6, detection would be a solved problem. You would just have to put the dial in the empty gap between 0.4 and 0.6 and the system would never make a mistake in either direction.
But the problem is that they do not sit in separate ranges. They never do. In 2023, Dalalah and colleagues found a substantial overlap between the score distributions of genuine and AI-generated academic writing. Later studies have found the same underlying overlap.
Now consider who ends up in that overlap.
Human writing lands in the machine range when it is clean and grammatically conservative. That includes a student writing carefully in a second language. It also includes technical and scientific prose, where uniformity is often a virtue, and short answers, which simply do not give the detector enough to work with. Jung and colleagues took that last problem seriously enough that in 2025 they built group-adaptive thresholds to stop short texts from being flagged at inflated rates.
Machine writing, on the other hand, lands in the human range when it is uneven, or when somebody with an understanding of human writing has edited it.
Once the two piles overlap, the trade-off is unavoidable. Wherever the dial sits, some human texts fall above it and get called machine-generated. And some machine texts fall below it and get identified as human. You cannot eliminate both errors simultaneously because both arise from the same overlapping region. Raise the dial to rescue the honest students, and you miss the AI texts sitting beside them. Lower it to catch those AI texts, and you will take the honest students with it.
And this is not a flaw in anybody’s product. No clever design and no ingenious engineering will ever escape it, because the trade-off belongs to the statistical overlap and not to the software.
Which mistake a school will tolerate
A school, or any educational institution for that matter, has to choose which of the two errors it is more willing to accept. But the problem is that these errors do not carry equal consequences. Missing an AI-written essay compromises an assessment. This is most likely not a big deal. Falsely accusing a student, however, can derail a degree and follow that student for years.
The consequence is that tools that flag innocent students get switched off because their vendors and the institutions that deploy them would otherwise risk being sued.
UCLA and the University of Pittsburgh disabled Turnitin’s AI detection early on. Vanderbilt did the same in August 2023. And OpenAI famously shut down its own detector a month earlier, having shipped a tool that caught only about 26 percent of AI text while falsely flagging about 9 percent of human text.
Flagging innocent students is not an option. And so the dial goes up. It has to, as long as language models continue to improve.
And that changes which numbers are really relevant. A vendor’s advertised accuracy comes from whatever threshold the vendor picked for its own benchmark. A school cannot use that same setting. It has to turn the dial up until the tool wrongly flags only, say, one honest essay in every hundred. And it then has to ask how much AI writing it still catches once the dial is that high.
In 2024, Tufts and colleagues argued that this was the only deployment metric worth reporting. They then tested popular detectors on text from models and subject areas the tools had not seen, using the kinds of prompts a curious student might try. Held to that one-in-a-hundred limit, some of those detectors caught nothing at all. Zero percent.
Now, you might object that these tools were simply not built to catch clever students. So let’s look at one that was. In 2025, Lekkala’s group tested a detector that had been trained specifically on AI text which someone had deliberately reworded to hide its origin. This system already knew what a disguised passage looks like. And yet, held to that same one-in-a-hundred limit, it caught only 48.8 percent of them.
So the safer the setting is for honest students, the less AI writing the detector is able to catch. And the less a dishonest student has to do to slip underneath it. The ones who still get caught at that setting are just the ones who did the least to hide.
Nobody writes down where the dial sits
Now, handing that decision back to the teacher sounds like the responsible thing to do, and in a sense, it is. Read the fine print on any of these products and you will find a version of the same warning. The score is not proof. It should not be the sole basis for an integrity finding. Every word of that is correct.
But look at what it actually does to the numbers. The vendor calculates its accuracy figure at a threshold the vendor selected, publishes it, and then tells you not to rely on it. Meanwhile, the threshold that decides whether a student gets an accusatory email is the one being set in your classroom, by you, and nobody has ever measured what that threshold does in practice.
And it never gets written down. Two instructors in the same department can look at the same 42 percent and draw the line in completely different places. And the same instructor can move it from one semester to the next without even noticing the change.
This is the part I find hardest to defend. A documented number becomes an undocumented disciplinary judgment, made at the end of a long grading day, about a student whose other writing the reader may barely remember. The trade-off between the two errors did not go away when the product stopped short of a verdict. It simply moved somewhere nobody keeps records.
The pile students can move
So far, I have treated the two piles as fixed. But they are not. Evasion creates a second problem, and it makes the trade-off even worse.
Many published benchmark results compare human writing against untouched model output. Someone pasted a prompt into ChatGPT, took the result, and scored it without changing a word. A student trying to avoid detection is highly unlikely to behave that way. They edit. They learn which prompts produce less machine-like prose, and some of them run the result through a tool built for exactly this. Every one of those moves drags the machine pile toward the human pile, and the effect is not marginal.
Krishna and colleagues built an eleven-billion-parameter paraphraser in 2023 and ran AI text through it. DetectGPT, held at a 1 percent false-positive rate, fell from catching 70.3 percent of that text to catching a mere 4.6 percent.
Liang’s group at Stanford needed no software at all. They asked ChatGPT to rewrite its own essays in more literary language, and detection fell from near 100 percent to roughly 13 percent. And Zhang and colleagues halved the performance of state-of-the-art detectors with nothing but changes to the user’s prompt.
Larger studies show the same effect. Perkins tested six detectors on 805 samples and found that ordinary manual editing reduced their accuracy by 17.4 percentage points. The researchers concluded the tools could not be recommended for integrity decisions. Similarly, Sun’s team tested thirteen detectors on over 280,000 pieces of real student coursework. A simple hybrid editing strategy let 88 percent of the AI content through.
I have run the small version of this myself, twice. In An Experiment in Language Laundering, I took an essay GPTZero rated 100 percent AI, pushed it through a commercial humanizer, and watched it come back rated 100 percent human. The whole workflow took under an hour. In Substack Can Now Scan Your Writing for AI, I published a post written end to end by Claude, laundered through the same tool, and Pangram cleared it as fully human with high confidence.
Moving the dial trades one error for the other, but at least the school controls that choice. Evasion widens the overlap itself.
Leave the dial where it is, and your honest students are as safe as they were. But more AI text now slips past. Turn it down to catch that text, and you start accusing honest students again. Either way, you end up worse off than before the students learned to evade it. The overlap grew, and no dial setting can shrink it back.
The advertised detection rate is therefore not just optimistic. It is an answer to a question about a world where nobody is trying to evade detection, published for a world where somebody always is.
Better detectors, same problem
The strongest counterarguments to my position come from researchers rather than vendors. Detectors do indeed improve. An independent 2025 audit by Jabarian and Imas at the University of Chicago found Pangram outperformed every commercial rival by a wide margin on a balanced corpus of nearly four thousand passages. Pangram represents genuine engineering progress. That is not disputed.
And the evasion problem also has known technical countermeasures. The same Krishna paper that broke DetectGPT with paraphrasing showed that searching a database of a provider’s past generations catches 80 to 97 percent of paraphrased text at a 1 percent false-positive rate. Kirchenbauer and colleagues showed statistical watermarks survive human paraphrasing once you have around 800 tokens to examine. And after documenting the 48.8 percent collapse, Lekkala’s group built an architecture that kept an 82.6 percent detection rate under the same attack.
But even granting all of it, the teacher still cannot tell who wrote the essay.
Retrieval searches the language model company’s own records, so it would only help if every provider ran the check. And a student using an open model on their own machine would leave no record with anyone, completely bypassing the system. Watermarking, meanwhile, is a tag the provider has to stamp into the text before a single word is generated, and most of them never did it. Where a watermark exists, it can be copied onto human writing and used to frame a student who cheated at nothing.
Neither of those approaches measures anything about the writing itself. They are bookkeeping. And they only work if there was an AI company involved in the first place, and if that company cooperated before the student ever opened the chat window.
Lekkala’s architecture is the one that actually reads the text, and it is genuinely better at what it does. But it was built against one very specific kind of evasion attack. No student is obliged to keep using that attack. And it never moved the two piles apart. It still has to draw its line through an overlap, which means it still has to choose which of the two errors it is more willing to accept.
Better detectors are possible, obviously. But the question is whether a better detector can solve the teacher’s problem. And as it turns out, it cannot.
Whose writing is the benchmark
Everything so far has come down to what people do. Schools raise the dial to protect honest students, and students learn to slip below it. But two limits sit underneath all of that, and neither has anything to do with how anybody behaves.
The first is the one most people have heard about. Sadasivan and colleagues proved that any detector’s performance is bounded by how different human and machine text are as distributions. And training a language model is, quite literally, the deliberate act of making those two distributions more alike. So as they converge, the best possible detector necessarily converges on a coin flip.
That is no one’s engineering failure. It is a ceiling that comes down a little further every time a better model ships.
Chakraborty and colleagues pushed back on this argument, though only in principle. As long as the two distributions differ anywhere at all, detection remains mathematically possible. You just need more samples. And that may well work if you are auditing thousands of submissions at once. But it does nothing for a teacher holding one essay by one student on a Tuesday afternoon.
The second limit is the one I find hardest to argue against, and it comes from Garland’s 2026 analysis. It turns out that the two-pile picture I have been using is already far too generous to the detector. It assumes the relevant comparison is between a text and human writing in general. But that is not the question a teacher is actually asking.
The question is whether this particular student wrote it. And a text-only detector has no model of that student whatsoever. It has never seen their other work, so it cannot tell a genuinely unusual sentence from a sentence that is merely unusual for them.
So it substitutes the population instead. It asks whether the text looks human in general. But human writing is enormously varied. Some real students write the way the machine writes. They did not cheat. Plain, conventionally structured prose is what many of them were taught to produce. For others, it is just how they write.
Any detector sensitive enough to be useful will flag some of these students. How many depends entirely on the overlap between genuine student writing and AI output. And that constraint would remain even if the models stopped improving tomorrow, because it comes from the diversity of the students themselves. Which is not a defect anyone should want to engineer away.
What replaces the scan
If the tool cannot answer the question, then the question has to be changed. The industry’s response, however, has so far been to move from detection to surveillance.
Turnitin Clarity records a drafting session for the instructor to replay, which means a student’s entire composing process now lives on a vendor’s servers. GPTZero’s extension performs a lighter form of the same monitoring inside Google Docs. But surveillance software does not become educational technology just because a school is the one that bought it.
I surveyed that market in Students Performing Human and came away convinced that process surveillance makes matters considerably worse. A student who thinks in long silences and then writes in bursts produces a suspicious-looking log. And so does a student using voice-to-text, because their words arrive in a block with no visible construction time, which is exactly the pattern the software treats as a paste event.
So watching the drafting simply creates more places for a false positive to happen. And it teaches students to perform composition for an observer instead of learning how to compose.
What has worked in my own classroom is simpler and slower. I ask students how they used AI before anything is graded, with no threat attached. Those conversations tell me more about students’ thinking than any scan ever has. But they only work when I drop the threat of an academic integrity violation.
And that approach also changes what I assess. If students can produce the finished artifact without doing the thinking the assignment was meant to develop, then the artifact is no longer enough evidence. Assessment has to move into the making of it.
Which means looking at the version history, the abandoned outline, the source journal noting where the student changed their mind, and the paragraph explaining what the model got wrong. What separates all of this from surveillance is consent and agency. A student documenting their own process is doing something a keystroke logger simply cannot do for them, because the documentation is itself an act of reflection.
The fourteen methods I collected in Fourteen AI-Proof Assessment Methods for the Age of Generative Intelligence all follow the same principle. A whiteboard defense works this way. So does an oral examination on a paper that the student wrote three weeks ago. Neither of them requires anyone to guess who typed what.
So it all comes down to this. A detector reports a number. Somebody then has to decide what that number is enough to justify. These days, that somebody is usually the teacher holding the report. And it is that decision, not the detector’s own setting, which determines which failure mode an educational institution is more willing to accept. It cannot eliminate both because some honest students and some language models will always produce writing that lands in the same range.
So, to protect the honest students, the dial has to go up. High enough that a student who spends just twenty minutes learning how to evade it passes beneath it easily. That student barely notices the line. The student who did nothing wrong is the one who has to answer for it.
Which is where the short version of this argument comes from. The more accurate AI detectors get, the easier it is for students to bypass them. Once a detector’s signal becomes a target, students learn to write around it. Goodhart’s Law, in a classroom.
That is the trade-off. It was never hidden. It was just never printed on the box.
The images in this article were generated with Nano Banana 2.
If you’d like to go further, the following NotebookLM-generated audio deep dive goes beyond the post into the broader research behind it, drawing on the sources and notes I gathered along the way. This is meant as a companion to the argument, offered as an optional extra rather than a summary of it.
P.S. I believe transparency builds the trust that AI detection systems fail to enforce. That’s why I’ve published an ethics and AI disclosure statement, which outlines how I integrate AI tools into my intellectual work.








This is the clearest version of the argument I've seen you make, and I think it actually understates its own strongest implication. The threshold/dial framing shows why no detector can eliminate both error types — that's airtight. But the piece that does the most work, for me, is buried two-thirds in: the Garland point that a text-only detector has no model of the student, only of "human writing in general." That's not a limitation that better engineering fixes. It's a category mismatch. The detector answers "did a machine statistically generate this string of text," and everyone downstream has been treating that as a stand-in for "did this student learn the material," which it never was, even before evasion entered the picture. A student who wrote every word herself but doesn't understand a thing she paraphrased will score as fully human and teach the grader nothing.
Which is really the question underneath all of this, and it predates AI entirely: how do you know if student X has learned Y? That was never something you could point an instrument at directly — it's always been inferred from proxies, and every proxy leaks somewhere. Essays leak because a finished artifact can (now, cheaply) be produced without comprehension. Oral defenses leak because fluency and confidence under pressure aren't the same thing as understanding, and they don't scale. The actual design problem was never "find the leak-proof proxy," because there isn't one — it was always about combining proxies whose failure modes don't overlap, so a student who fakes one channel gets exposed on another. A cold oral defense weeks after submission, without notes, doesn't test the same thing an essay does; it tests whether the material lives in the student's head at all. That's not a superior detector. It's two bad instruments stacked so their blind spots don't coincide.
Which is maybe the real scandal in the detection industry, sharper than any accuracy number: it let institutions forget the tool was never measuring the thing they actually cared about.