Why Some AI Drafts Resist Editing
On text that looks polishable but really isn't.
This post is going out early as a thank-you to paid subscribers. The full piece opens up to free subscribers next Tuesday. I’m grateful you’re here, and grateful for your commitment to rethinking what education can be. Your support is what keeps this work going.
A few months ago, I published “A Year of AI-Assisted Writing,” a piece describing my AI-assisted writing process. It was a follow-up to my ethics statement, and it laid out in some detail how a blog post moves from an idea in my head to a finished piece on The Augmented Educator: an AI-assisted draft followed by heavy, iterative human editing. I wrote it because many Substack authors appear to work the same way without ever saying so, and the approach still feels controversial enough that I wanted mine disclosed properly.
What I did not talk about in that piece is that not every idea that starts in my head makes it onto The Augmented Educator. Sometimes the topic turns out to be less interesting than I thought. Sometimes it ends up a tad too technical. I have a drawer full of essay concepts about cybersecurity in the AI age, but most of them would not appeal to the audience of this Substack.
More often than not, however, an idea dies because the initial AI-generated draft resists human editing at a level I did not expect when I started using this workflow.
There are drafts that simply fall into place, where the editing feels natural and a few iterations produce a consistent piece that flows. And there are drafts that look polished on the surface but fall apart the moment you start cleaning things up. It is not unheard of for me to reach a point where I simply give up. A point where the editing effort gets me nowhere, where every attempt to fix one problem only surfaces two new ones.
People sometimes reach for the saying that “you cannot polish a turd,” and for a long time that was my private shorthand too. But the saying does not quite describe the problem. A turd announces itself. Nobody picks one up expecting to polish it.
The drafts I am talking about look clean, read well, and pass every quick inspection, and the trouble only shows once the polishing has begun and hours are already spent. The draft was never bad in any traditional sense. It was just not a starting point from which my iterative workflow had any chance of converging on a piece I would consider fit for my readers. The saying, if anything, gets the situation backwards. The problem is that these drafts look eminently polishable. They just aren’t.
I have always wondered why that is, and how to make sure every draft I generate can become a publishable essay rather than a mess of never-ending edits.
I suspected the answer might also explain why many professional writers — people who can write perfectly well without assistance — so often struggle with editing AI output. To be clear, I am not suggesting they should trade their tried-and-true approach for an AI workflow. But if there were more clarity about why AI text sometimes resists human editing, it could open up better pathways for teaching AI literacy to professionals who have tried these tools and found them wanting.
So in today’s essay, I want to dig into the question of why some AI-generated texts appear polished on the surface yet are fundamentally flawed to the point of being uneditable, while others just work. And I want to explore how to raise the odds that a prompt produces an internally consistent draft in the first place.
When fluency is a disguise
The first reason an editor might struggle with an AI draft lies in how the text is made. Large language models generate prose by predicting the next token, one after another, in whatever sequence is statistically most likely given the training data. This mechanism is excellent at reproducing the surface of authoritative writing. Computational linguists have a name for the result: deceptive fluency.
A deceptively fluent text is grammatically clean, smoothly connected, and formatted exactly as its genre demands, yet hollow underneath. In academic and educational contexts, this produces what some researchers call the fluency fallacy: our tendency to mistake coherent academic language for genuine understanding.
Reviewers of AI-drafted literature reviews report the pattern again and again: generic explanations, repetitive sentence structures, weak critical analysis, and conclusions broad enough to fit any paper. The model can summarize ten studies in seconds. It almost never notices the tension between two of them.
An editor who sits down to polish a draft is operating on an assumption: that the draft has a sound foundation, and that the remaining work is surface work. Fix the phrasing, add domain expertise, sharpen the examples. With a deceptively fluent draft, that assumption is false. The correct punctuation and the smooth transitions sit on top of an argument made of filler, statistically probable sentences arranged in the shape of reasoning.
And the surface is stubborn. Because the model’s prose is so tightly woven at the sentence level, inserting one genuinely analytical thought tends to break the flow of everything around it.
You fix a paragraph and the section stops hanging together. You fix the section and the essay’s through-line snaps. The draft gets abandoned, in the end, because its artificial coherence cannot carry the weight of actual reasoning. Untangling the machine’s surface logic costs more than articulating the thought from scratch would have.
The inspector on the conveyor belt
To understand why tearing down and rebuilding a draft is so exhausting, it helps to look at what post-editing does to the writer’s brain. Traditional models of writing describe a cycle of planning, translating ideas into text, and reviewing. An AI-first workflow reshuffles this cycle. The writer stops being a creator and becomes a reviewer of someone else’s output, and that shift changes the cognitive economics of the whole task.
Cognitive load theory sorts mental effort into three kinds: intrinsic load, the inherent difficulty of the task; extraneous load, the wasted effort imposed by bad tools and friction; and germane load, the productive effort that builds understanding.
The promise of AI drafting is that it absorbs the intrinsic load of getting ideas into words, freeing the writer’s working memory for higher-order thinking. For some writers, this promise holds. Studies of second-language learners, for instance, find that AI assistance genuinely lifts the burden of grammatical mechanics, and the learners notice it.
For an experienced writer editing a full draft, something stranger happens. The typing effort drops, but the effort of evaluation goes up, and it goes up a lot.
Reading AI output is not like reading a colleague’s draft. With a colleague, you can trust that there is an intent behind every paragraph, a lived experience, a mental model you share. With a model, you can trust none of that. Every claim might be hallucinated, every transition might be papering over a gap, every confident sentence has to be checked. Researchers developing cognitive load scales for AI-assisted writing have decomposed this into distinct factors — prompt management and critical evaluation — that did not exist in the older models.
The practical consequence is a phenomenon that practitioners have taken to calling AI fatigue or review fatigue. Judging whether a generated paragraph matches your intent requires a stream of small verdicts, hundreds of them an hour. Hold the machine’s logic in working memory, compare it against your own knowledge, spot the discrepancy, plan the fix, repeat.
Writing from scratch is a proactive state in which the writer builds an arc of coherence at their own pace. Post-editing puts the same writer in the position of a quality inspector on a conveyor belt that never stops.
There is a bitter twist at the end of this. Fatigue degrades exactly the faculty the inspector needs most, which is judgment. A worn-down editor starts trusting the machine too readily. The literature calls this automation bias, and it means the drafts most likely to slip through unfixed are the ones that arrived when the editor had nothing left.
The forty-percent line
The feeling that a draft is unpolishable is not just a mood. It can be measured, and an entire industry has been measuring it for decades. Machine translation post-editing has long needed to know when correcting a machine’s output stops being cheaper than translating from scratch, and its metrics transfer surprisingly well to AI-assisted writing.





