This is the clearest version of the argument I've seen you make, and I think it actually understates its own strongest implication. The threshold/dial framing shows why no detector can eliminate both error types — that's airtight. But the piece that does the most work, for me, is buried two-thirds in: the Garland point that a text-only detector has no model of the student, only of "human writing in general." That's not a limitation that better engineering fixes. It's a category mismatch. The detector answers "did a machine statistically generate this string of text," and everyone downstream has been treating that as a stand-in for "did this student learn the material," which it never was, even before evasion entered the picture. A student who wrote every word herself but doesn't understand a thing she paraphrased will score as fully human and teach the grader nothing.
Which is really the question underneath all of this, and it predates AI entirely: how do you know if student X has learned Y? That was never something you could point an instrument at directly — it's always been inferred from proxies, and every proxy leaks somewhere. Essays leak because a finished artifact can (now, cheaply) be produced without comprehension. Oral defenses leak because fluency and confidence under pressure aren't the same thing as understanding, and they don't scale. The actual design problem was never "find the leak-proof proxy," because there isn't one — it was always about combining proxies whose failure modes don't overlap, so a student who fakes one channel gets exposed on another. A cold oral defense weeks after submission, without notes, doesn't test the same thing an essay does; it tests whether the material lives in the student's head at all. That's not a superior detector. It's two bad instruments stacked so their blind spots don't coincide.
Which is maybe the real scandal in the detection industry, sharper than any accuracy number: it let institutions forget the tool was never measuring the thing they actually cared about.
Correct. As educators we need to ask the question if student X has learned Y and the only way to do that, in my opinion, is to actually interact with the student and to not just judge the output.
This is the clearest version of the argument I've seen you make, and I think it actually understates its own strongest implication. The threshold/dial framing shows why no detector can eliminate both error types — that's airtight. But the piece that does the most work, for me, is buried two-thirds in: the Garland point that a text-only detector has no model of the student, only of "human writing in general." That's not a limitation that better engineering fixes. It's a category mismatch. The detector answers "did a machine statistically generate this string of text," and everyone downstream has been treating that as a stand-in for "did this student learn the material," which it never was, even before evasion entered the picture. A student who wrote every word herself but doesn't understand a thing she paraphrased will score as fully human and teach the grader nothing.
Which is really the question underneath all of this, and it predates AI entirely: how do you know if student X has learned Y? That was never something you could point an instrument at directly — it's always been inferred from proxies, and every proxy leaks somewhere. Essays leak because a finished artifact can (now, cheaply) be produced without comprehension. Oral defenses leak because fluency and confidence under pressure aren't the same thing as understanding, and they don't scale. The actual design problem was never "find the leak-proof proxy," because there isn't one — it was always about combining proxies whose failure modes don't overlap, so a student who fakes one channel gets exposed on another. A cold oral defense weeks after submission, without notes, doesn't test the same thing an essay does; it tests whether the material lives in the student's head at all. That's not a superior detector. It's two bad instruments stacked so their blind spots don't coincide.
Which is maybe the real scandal in the detection industry, sharper than any accuracy number: it let institutions forget the tool was never measuring the thing they actually cared about.
Correct. As educators we need to ask the question if student X has learned Y and the only way to do that, in my opinion, is to actually interact with the student and to not just judge the output.