Why Some AI Drafts Resist Editing
On text that looks polishable but really isn't.
A few months ago, I published “A Year of AI-Assisted Writing,” a piece describing my AI-assisted writing process. It was a follow-up to my ethics statement, and it laid out in some detail how a blog post moves from an idea in my head to a finished piece on The Augmented Educator: an AI-assisted draft followed by heavy, iterative human editing. I wrote it because many Substack authors appear to work the same way without ever saying so, and the approach still feels controversial enough that I wanted mine disclosed properly.
What I did not talk about in that piece is that not every idea that starts in my head makes it onto The Augmented Educator. Sometimes the topic turns out to be less interesting than I thought. Sometimes it ends up a tad too technical. I have a drawer full of essay concepts about cybersecurity in the AI age, but most of them would not appeal to the audience of this Substack.
More often than not, however, an idea dies because the initial AI-generated draft resists human editing at a level I did not expect when I started using this workflow.
There are drafts that simply fall into place, where the editing feels natural and a few iterations produce a consistent piece that flows. And there are drafts that look polished on the surface but fall apart the moment you start cleaning things up. It is not unheard of for me to reach a point where I simply give up. A point where the editing effort gets me nowhere, where every attempt to fix one problem only surfaces two new ones.
People sometimes reach for the saying that “you cannot polish a turd,” and for a long time that was my private shorthand too. But the saying does not quite describe the problem. A turd announces itself. Nobody picks one up expecting to polish it.
The drafts I am talking about look clean, read well, and pass every quick inspection, and the trouble only shows once the polishing has begun and hours are already spent. The draft was never bad in any traditional sense. It was just not a starting point from which my iterative workflow had any chance of converging on a piece I would consider fit for my readers. The saying, if anything, gets the situation backwards. The problem is that these drafts look eminently polishable. They just aren’t.
I have always wondered why that is, and how to make sure every draft I generate can become a publishable essay rather than a mess of never-ending edits.
I suspected the answer might also explain why many professional writers — people who can write perfectly well without assistance — so often struggle with editing AI output. To be clear, I am not suggesting they should trade their tried-and-true approach for an AI workflow. But if there were more clarity about why AI text sometimes resists human editing, it could open up better pathways for teaching AI literacy to professionals who have tried these tools and found them wanting.
So in today’s essay, I want to dig into the question of why some AI-generated texts appear polished on the surface yet are fundamentally flawed to the point of being uneditable, while others just work. And I want to explore how to raise the odds that a prompt produces an internally consistent draft in the first place.
When fluency is a disguise
The first reason an editor might struggle with an AI draft lies in how the text is made. Large language models generate prose by predicting the next token, one after another, in whatever sequence is statistically most likely given the training data. This mechanism is excellent at reproducing the surface of authoritative writing. Computational linguists have a name for the result: deceptive fluency.
A deceptively fluent text is grammatically clean, smoothly connected, and formatted exactly as its genre demands, yet hollow underneath. In academic and educational contexts, this produces what some researchers call the fluency fallacy: our tendency to mistake coherent academic language for genuine understanding.
Reviewers of AI-drafted literature reviews report the pattern again and again: generic explanations, repetitive sentence structures, weak critical analysis, and conclusions broad enough to fit any paper. The model can summarize ten studies in seconds. It almost never notices the tension between two of them.
An editor who sits down to polish a draft is operating on an assumption: that the draft has a sound foundation, and that the remaining work is surface work. Fix the phrasing, add domain expertise, sharpen the examples. With a deceptively fluent draft, that assumption is false. The correct punctuation and the smooth transitions sit on top of an argument made of filler, statistically probable sentences arranged in the shape of reasoning.
And the surface is stubborn. Because the model’s prose is so tightly woven at the sentence level, inserting one genuinely analytical thought tends to break the flow of everything around it.
You fix a paragraph and the section stops hanging together. You fix the section and the essay’s through-line snaps. The draft gets abandoned, in the end, because its artificial coherence cannot carry the weight of actual reasoning. Untangling the machine’s surface logic costs more than articulating the thought from scratch would have.
The inspector on the conveyor belt
To understand why tearing down and rebuilding a draft is so exhausting, it helps to look at what post-editing does to the writer’s brain. Traditional models of writing describe a cycle of planning, translating ideas into text, and reviewing. An AI-first workflow reshuffles this cycle. The writer stops being a creator and becomes a reviewer of someone else’s output, and that shift changes the cognitive economics of the whole task.
Cognitive load theory sorts mental effort into three kinds: intrinsic load, the inherent difficulty of the task; extraneous load, the wasted effort imposed by bad tools and friction; and germane load, the productive effort that builds understanding.
The promise of AI drafting is that it absorbs the intrinsic load of getting ideas into words, freeing the writer’s working memory for higher-order thinking. For some writers, this promise holds. Studies of second-language learners, for instance, find that AI assistance genuinely lifts the burden of grammatical mechanics, and the learners notice it.
For an experienced writer editing a full draft, something stranger happens. The typing effort drops, but the effort of evaluation goes up, and it goes up a lot.
Reading AI output is not like reading a colleague’s draft. With a colleague, you can trust that there is an intent behind every paragraph, a lived experience, a mental model you share. With a model, you can trust none of that. Every claim might be hallucinated, every transition might be papering over a gap, every confident sentence has to be checked. Researchers developing cognitive load scales for AI-assisted writing have decomposed this into distinct factors — prompt management and critical evaluation — that did not exist in the older models.
The practical consequence is a phenomenon that practitioners have taken to calling AI fatigue or review fatigue. Judging whether a generated paragraph matches your intent requires a stream of small verdicts, hundreds of them an hour. Hold the machine’s logic in working memory, compare it against your own knowledge, spot the discrepancy, plan the fix, repeat.
Writing from scratch is a proactive state in which the writer builds an arc of coherence at their own pace. Post-editing puts the same writer in the position of a quality inspector on a conveyor belt that never stops.
There is a bitter twist at the end of this. Fatigue degrades exactly the faculty the inspector needs most, which is judgment. A worn-down editor starts trusting the machine too readily. The literature calls this automation bias, and it means the drafts most likely to slip through unfixed are the ones that arrived when the editor had nothing left.
The forty-percent line
The feeling that a draft is unpolishable is not just a mood. It can be measured, and an entire industry has been measuring it for decades. Machine translation post-editing has long needed to know when correcting a machine’s output stops being cheaper than translating from scratch, and its metrics transfer surprisingly well to AI-assisted writing.
The workhorse metrics are “Post-Edit Distance” and “Translation Edit Rate.” Both count the minimal operations, the insertions, deletions, substitutions, and shifts, needed to turn a machine draft into the approved final version.
The research on these metrics points to a clear threshold. When edits touch roughly 40 percent of a machine-generated text, the effort of post-editing overtakes the effort of writing from scratch. Past that line, the draft is uneconomical to save. Not as a matter of taste. As a matter of arithmetic.
Keystroke counts do not tell the entire story, though, because technical effort and cognitive effort are not the same thing. A single semantic flaw might take ten keystrokes to fix and twenty minutes to find. Translation researchers capture this with the pause-to-word ratio. Cognitively demanding output produces clusters of brief pauses in which the editor is reading, re-reading, and deciding how to intervene. A draft can score well on edit distance and still be a cognitive swamp.
This is, I think, the empirical shape of the moment I described in the introduction, the moment of giving up. The writer is holding a disjointed machine narrative in working memory while simultaneously trying to plan the coherent structure that should replace it, and that double duty exceeds what working memory can do.
Somewhere, consciously or not, the writer runs the numbers and concludes that the honest estimate is past the threshold. The rational move is to stop editing and start over.
The first tracks become the rut
So far the problems have lived in the draft. The next one lives in us. Cognitive scientists and design researchers call it design fixation, or in its behavioral-economics form, anchoring: the unconscious adherence to an initial concept that narrows the search space and blocks better solutions.
Before generative AI, writers began with low-fidelity material. Outlines, mind maps, scribbled notes, or placeholder text. The roughness was a feature. A sketch tells your brain that everything is still negotiable.
An LLM skips the sketch entirely and hands you finished-looking prose, and that high-fidelity surface sends a quiet signal that the text is settled, mature, in need of nothing more than touch-ups. The first tracks laid down in a problem space easily become the rut a person follows. When the machine lays the first tracks, the writer ends up playing on the machine’s terms.
Empirical work shows this. Studies comparing human-only, LLM-only, and collaborative writing find an asymmetry: models score high on structural coherence and grammatical execution, while humans keep a clear edge in originality and divergent thinking on demanding creative tasks.
But when humans start from a machine draft, they often fall into what one research team named the collaboration trap. The draft anchors them. They mimic its complexity without adding quality of their own, and because the model’s output gravitates toward the statistical median of its training data, the anchored result drifts toward the generic.
This is the second way a draft can defeat its editor, and it is the more insidious one. The text executes a mediocre idea flawlessly. It is too well-written to throw away without a pang, and too generic to matter to anyone. The editor polishes the surface while the underlying idea stays anchored to the machine’s uninspired baseline.
Ghostwriter or sounding board
Whether a draft anchors you depends a great deal on who you are. A study published in Management Science examined this directly, assigning expert and novice writers to different modes of collaboration with an LLM, and its findings map neatly onto the frustrations I hear from professional writers.
For novices, an AI draft raises the floor. Someone without a developed internal model of what the final text should look like experiences little friction when editing machine output, because the machine’s generic fluency is a genuine upgrade on their unassisted work. The draft overcomes the blank page and organizes scattered thoughts. The post-editing feels rewarding because it is.
For experts, the dynamic inverts. An expert arrives with a developed mental model, a distinctive voice, and specific intentions, and when the model plays ghostwriter, generating the body of the text, every one of those assets becomes a source of friction. The expert spends the session reconciling a nuanced internal vision with a statistically average external draft, deleting, restructuring, fighting the machine’s cadence.
The same study found the fix in a change of role. When experts kept the initial generation for themselves and used the model as a sounding board — a critic that flags logical gaps, suggests local rephrasing, and checks tone — the collaboration paid off. The expert maintains the architecture. The machine inspects it.
I find this result clarifying because it dissolves an apparent paradox. The writers most likely to dismiss AI assistance as useless are often the most skilled, and the standard explanation is stubbornness. The better explanation is that they have been handed the one collaboration mode that is worst for them.
The voice that will not wash out
There is one more failure mode, and it can sink a draft that is structurally sound and factually clean. Style. A writer’s voice is a set of deliberate linguistic habits that signal identity and audience alignment, and for anyone publishing under their own name, on a platform where the voice is the brand, it is not negotiable.
AI prose has a voice of its own, and readers have learned to hear it. The predictable cadence, the stock transitions, the sanitized diplomatic tone that flattens every idiosyncrasy.
The interesting question is whether editing can wash it out, and a pre-registered study by Baumler and colleagues, presented at ACL 2026, set out to measure exactly that. Using embedding-based authorship representations that capture stylistic fingerprints at a fine grain, the researchers compared raw LLM text, human-post-edited LLM text, and text that the same authors wrote from scratch.
The result is sobering. Post-editing moves a draft toward the author’s natural voice, but only about a quarter of the way. The edited drafts remained measurably closer to the raw machine text than to the author’s own unassisted writing, and their stylistic range was narrower than the human baseline. The machine’s fingerprints survive the polish.
The study found a perception gap, too. Editors reported high satisfaction with their polished drafts and perceived them as authentically their own, even as the metrics detected the residue.
I recognize this from my own workflow, and it worries me more than the residue itself. Over time, writers who lean heavily on generation report a creeping sense of voice loss, the recognition that their output has become competent and sterile. By the time you hear it yourself, your readers will have been hearing it for a while.
Getting a draft you can actually edit
Everything above explains the failures. It also, read in reverse, tells you how to succeed, and this is where I want to spend the rest of the essay. The common thread through deceptive fluency, anchoring, and the style gap is premature prose. Every one of these traps is sprung by finished-looking text arriving before the human has settled the structure and the intent. The remedies all amount to the same approach: keep the machine away from prose until the structure of the piece is yours.
The first and most consequential strategy is to separate structure from style. The dominant habit of one-shot prompting, asking for a complete draft in a single request, guarantees an anchor and a style gap. The alternative is to collaborate on the skeleton.
Some researchers describe this as working with structural beats: the writer and the model iterate on a detailed sequence of arguments, evidence, and turns, presented as an outline rather than as flowing text. Because an outline is visibly unfinished, it does not trigger the settledness signal that polished prose does.
Everything still reads as negotiable because it is. Only when the beats are fixed does one write sentences. And the sentences can then be written by the human author, or generated in small, tightly constrained modules.
The second strategy formalizes the first into a pipeline. Research systems such as Papers-to-Posts show a decoupled plan-draft-revise loop: the model proposes a modular plan, typically bullet points extracted from source material; the human selects, deletes, and reorders those bullets; and only the approved plan gets expanded into text.
The draft that emerges is built on a foundation the human has already inspected. In my experience, this is the single biggest lever for editability. A draft grown from a plan I have pruned almost always converges. A draft conjured from a one-line prompt is a coin flip.
The third strategy concerns the interface. Chat windows encourage a hands-off posture: the text appears in the machine’s space, whole, and you react to it.
Experimental systems point in a different direction. Tools like AnchoredAI attach the model’s suggestions to specific spans of the writer’s own text, which measurably strengthens the writer’s sense of ownership and keeps the human’s tracks dominant on the page. Canvas-style prompting environments such as PromptCanvas break the linear chat into rearrangeable widgets, so that iterating on an idea stops feeling like scrolling through a transcript.
This idea generalizes even if you never touch these tools: work in your document, pull the machine in for localized tasks, and be suspicious of any workflow in which the model’s window is the primary one.
The fourth strategy is the strangest and, in my practice, the second most valuable. Design researchers have explored what they call machine unlearning in creative collaboration: deliberately constraining the model away from its statistical comfort zone.
In practice, this means prompting with hard exclusions. Forbid the stock transitions. Ban the words and constructions you have learned to recognize as machine tells. Rule out the standard essay shapes. Constraints like these force the model off its probabilistic median, and the output that comes back is rougher, odder, and less finished-looking. That roughness is the point.
A slightly raw draft invites the editor in. A gleaming one warns the editor off, and we have seen where that leads.
Notice what the four strategies have in common. Each one delays finished prose until the human has signed off on what the prose will say. The order of operations is the quality control.
What I have changed in my own workflow
I started this essay with a puzzle from my own practice: why some drafts converge under editing and others eat every hour I give them. Having spent real time with this research, I can now name what the good drafts had in common, and it is not luck.
When I look at the pieces that came together, I find strict adherence to the plan-draft-revise approach. The structure existed, and I had approved it, before any prose was generated. And I find adherence to the use of unlearning constraints, the explicit list of forbidden words, stock moves, and machine cadences that I now attach to every drafting prompt.
When I look at the pieces I abandoned, I almost always find a shortcut. A one-shot draft I asked for because I was tired, or a plan I skimmed instead of pruned. The drafts that defeated me were commissioned, not encountered.
If you write with AI assistance, or teach people who do, I think this reframing is the useful takeaway. Editability is not a property you discover in a draft after the fact. It is a property you build in before the first sentence is generated, through the order of operations you impose on the machine. The research on thresholds suggests there is little middle ground: a draft either starts close enough to your intent that editing converges, or it does not, and no amount of stubbornness at the keyboard moves it across that line.
Which is why I have stopped asking how to edit better. The drafts that could be saved never needed much saving. The rest were lost before I typed a word of my own.
The images in this article were generated with Nano Banana 2.
If you’d like to go further, the following NotebookLM-generated audio deep dive goes beyond the post into the broader research behind it, drawing on the sources and notes I gathered along the way. This is meant as a companion to the argument, offered as an optional extra rather than a summary of it.
P.S. I believe transparency builds the trust that AI detection systems fail to enforce. That’s why I’ve published an ethics and AI disclosure statement, which outlines how I integrate AI tools into my intellectual work.








