Can Kimi K3 Write?
Three frontier models wrote the same essay. Four AI judges picked a winner.
On July 16, Moonshot AI released Kimi K3, which the company describes as the first open-weight model in the three-trillion-parameter class. As I am writing this, the weights themselves are not yet out. Moonshot has said all 2.8 trillion of them will be published on July 27 under a modified MIT license.
I should note that “open weights” does not mean “local execution” in this case. In its native format, the model requires roughly a terabyte and a half of video memory. That is before you count anything else that has to sit in memory alongside the weights. Nobody will be able to run this at home.
But what open weights buy is open competition: once the files are public, third-party hosts will be able to serve K3 without asking anyone’s permission, at a launch price low enough that the proprietary labs will have to answer it. For the first time, the open-weight ecosystem is breathing down the necks of Anthropic’s Claude Fable 5 and OpenAI’s ChatGPT 5.6 Sol.
That industry story is interesting in itself. But the question I want to address in this newsletter is more personal. Can the thing write? I mean, not benchmark-write. Write under constraints, with a voice a human editor would want to work with.
So in today’s post, I want to walk through an experiment I ran to find out exactly that: one research brief, one style guide, three models drafting, four judging blind, and a closing round of guess-who-wrote-what that went almost too well.
One brief, no scaffolding
Regular readers know the drafting workflow I usually follow. I have described it in an earlier post.
I start by sending an essay idea into Gemini 3.5 Deep Research to develop a sourced brief. The brief then goes to a language model, usually the latest Claude model, along with The Augmented Educator style guide. The model returns a rough draft. Only then does the actual writing start. I use an iterative process and rewrite until it sounds like me.
The essay idea I used for this experiment came from a YouTube video in which the creator made the following deceptively logical claim: “In art, the effort does not matter. The art itself matters.” His thesis was that talented artists will thrive with AI tools, whereas untalented artists will fall behind regardless of AI use.
I felt there might be some historical context to unpack here, since this is likely not the first time this claim was made. It therefore seemed like a good seed for an essay on what happens to art education when machines absorb the effort. Perfect for The Augmented Educator.
For the purpose of this experiment, I applied one major deviation from my routine. Usually, the prompt I use for drafting includes detailed instructions about story angle and structure. Here, I withheld all of it. The models got the brief and the guide, nothing more. I wanted to see and evaluate their raw judgment and writing skills, and not my own scaffolding mirrored back at me.
The result was three drafts by three contenders:
Claude Fable 5 (run on its Max setting) wrote “Take the Hand from the Picture.”
ChatGPT 5.6 Sol (run on its Pro setting) wrote “The Art Does Not Come With a Timesheet.”
Kimi K3 (run on its Max setting) wrote “The Panel with Three Lines.”
The essays themselves were not really remarkable, and to be honest, none of them would make it onto this Substack without very heavy editing, if at all. I am attaching them here only for reference in case someone wants to check them out. But regardless, I felt they were good enough for the experiment.
A three-to-one landslide
I handed the three unlabeled essays, plus the style guide, to four AI judges for a blind review. These judges were fresh instances of the three author models and Gemini 3.5 Pro, the model that generated the brief. Each judge scored every essay from 1 to 10 across three categories: style-guide adherence, prose quality, and argument and storytelling. Each model was also asked to pick exactly one essay to publish.
In the following table, each figure is one judge’s three category scores averaged into a single mark out of 10 for that essay. And the bottom row averages those marks across all four judges.
If you look at these numbers, one question pops up immediately. Why did Fable’s draft win so clearly? Two main reasons came up in three of the four verdicts.
The first was the model’s sheer discipline. It hit the guide’s fussiest targets, and it hit them visibly: wherever a list wanted three items, Fable wrote four, dodging the banned rule of three. It was also the only draft that dug past the obvious material in the brief and used the specialist studies buried further down, which the other two left untouched.
The second reason was intellectual. Fable’s was the only draft that made the logical collapse of the YouTuber’s quote its thesis rather than a passing correction. Because if talent is a fixed quantity that tools merely expose, as the claim indicates, then teaching art is a pointless exercise. Fable noticed this inconsistency in the brief’s analysis and built its entire narrative around it.
Sol dissented on both counts. It was the only judge that marked Fable’s draft down on adherence, faulting the long paragraphs and a stack of balanced contrasts. And it thought its own draft handled the talent question more cleanly.
The other texts split the judges. Sol’s “Timesheet” drew genuine praise for separating what is worth encountering as art from what is worth assigning as education. But its bolded imperative takeaways read to Fable like a faculty memo, and to Kimi like a workshop handout drifting toward the generic “5 ways to...” article the guide warns against. Fable also found its clipped, uniform rhythm the most machine-like in the pool.
Kimi’s “Panel” split the room differently. Every judge praised the liveliness of its sentences. Two ranked the draft second on points, but none of them recommended publishing it first. And Sol liked the individual sentences but disliked the voice they added up to, calling it prosecutorial rather than provocative and short on generosity toward students.
Before I continue, I need to add a quick note on bias, because some readers are probably already typing in the comment section. Language models grading language models is somewhat of a circular exercise, and self-preference is a documented failure mode. Sure enough, the two proprietary author models, Fable and Sol, each picked their own work. Gemini, with no skin in the game, sided firmly with the majority.
Guess who wrote what
After scoring, I asked each judge to guess which essay was written by which model. Three of the four guessed perfectly. Each model, it turns out, wrote with some recognizable habits:
Fable followed instructions to the letter, down to the intentional four-item lists.
Sol leaned on a rigid structure, characterized by short, symmetrical sentences and takeaways formatted as bolded lists.
Kimi K3 wrote with a casual punch and put momentum above the fine print. This is exactly how the banned rule of three slipped back into its essay.
Kimi identified “Panel” as its own work because it “treats the guide as a vibe rather than a spec.” That is a sharp piece of self-recognition from a model that had ranked that same essay dead last a few minutes earlier. Only Gemini stumbled. It described the habits accurately but filed two of them under the wrong names, pinning Sol’s bolded lists on Kimi and Kimi’s rule-of-three slips on Sol.
The judges also proactively disclosed the limits of the exercise. Sol footnoted an arXiv paper, warning me that model attribution remains an unsolved research problem. Kimi insisted upfront that models have no privileged ability to recognize their own text. And Fable volunteered a disclosure that, as a Claude model, it might be flattering its own family.
So can Kimi actually write?
Now to the core question. Can Kimi K3 write at the level of the leading flagship models?
Let’s start with the good news, which is the voice. Two judges called Kimi’s hook the strongest of the three, and its paragraphs move better than anything else in the pool. Nothing in it sounds like the beige, press-release text we so often associate with machine drafts.
As for the bad news, Kimi’s draft broke the guide’s most explicit ban, and broke it repeatedly. The rule of three is back on nearly every page, although that would be easy to fix in post-editing.
The larger problem was the logic. The essay endorsed the YouTuber’s claim, then pivoted to a lesson plan anyway. That is the contradiction Fable’s draft made its thesis, and the one Sol’s draft avoided by refusing the fixed-talent premise. Kimi carried it to the final paragraph without noticing.
As I have written before, I treat a model draft like a delivery of clay. The material is real, but the shape isn’t mine yet. That clay-carving stage is where the actual writing happens for me, and it is why Kimi’s logical lapses are more consequential to me than its stray triads.
What educators should take away
This admittedly imperfect experiment suggests that Kimi K3 can write close to the level of the leading foundational models. Its draft finished essentially level with Sol’s and well behind Fable’s, which is still a startling place for an open-weight model to land. However, I will probably stay with Claude Fable 5 for my drafting workflow, at least for as long as Fable is included in my Claude subscription.
That could change, though, because open weights alter the financial arithmetic underneath the whole comparison. Fable on its Max setting costs real money. By contrast, Moonshot launched K3 at three dollars per million input tokens and fifteen per million output. And once the files are public, no single vendor controls that meter. It is cheaper, swappable, inspectable, and impossible to un-release.
For an educator or a small publication choosing a drafting engine on a budget, K3 is the first open model I would put into a writing workflow without hesitation.
Beyond the budget, the real surprise was the guessing game. If three out of four models can identify each other from a single sample of prose, those tells are learnable, and not only by machines. A teacher who reads enough model output will start recognizing these signatures without specialized detection software.
But I want to be careful here, because the judging models had an easier task than a teacher does. They were choosing from three named candidates, all working from the same brief and the same guide, with no human revision in between. Spotting an edited student paper with no candidate list and no control text is a much harder problem.
Still, the direction is clear. Reading model output closely is becoming part of the job, and that is a skill built by exposure.
So, can Kimi K3 write? It can. It writes well enough to sit just one tier below Fable at a fraction of the cost, and open weights mean the market will decide how much that remaining gap is worth, at least until the next release forces us to recalculate.
Every model signs its drafts. The rewrite is how I sign mine.
The images in this article were generated with Nano Banana 2.
P.S. I believe transparency builds the trust that AI detection systems fail to enforce. That’s why I’ve published an ethics and AI disclosure statement, which outlines how I integrate AI tools into my intellectual work.






