A Workbench Is Not a Soul
What Claude's newly discovered "conscious access" means for the classroom
On July 6, Anthropic published a research paper with the unglamorous title “Verbalizable Representations Form a Global Workspace in Language Models.” Its sixteen authors report that Claude, the company’s language model, maintains a small, privileged set of internal representations. These function as a silent working memory where the model holds concepts and reasons with them before a single word appears on screen.
Nobody built this structure. It emerged on its own during training. And because it mirrors a leading neuroscientific theory of how humans consciously access information, the paper uses the term “conscious access” throughout.
You can probably guess what happened next. Within a day, social media feeds were filled with confident declarations that Claude is conscious. One widely shared headline announced that Anthropic now thinks Claude has a soul. And screenshots of the paper’s odder findings circulated with captions about machine sentience and inner lives.
What I find most irritating is how this completely misrepresents the paper. The researchers state, plainly, that they take no position on whether Claude has subjective experience. In their work, the term “conscious” has a very specific technical definition, and the difference between that and the common usage is precisely where the public discussion went off track.
The research itself, though, is substantial. And for educators, it might be more important than almost anything published on AI this year. It changes what we need to teach students about these systems, because it gives us, for the first time, a real way to look inside one.
So in this essay, I want to walk through what the paper shows, the claims it carefully declines to make, and how all of this relates to the classroom.
This leads me to the workbench metaphor used in my title. If you take the paper’s “global workspace” literally, you get the image of a bench in the middle of a large workshop. A bench can only hold a few parts at a time. The items laid out on it are accessible for anyone in the shop to use. And the bench has no feelings about the work it supports.
Keep that bench in mind. Most of what follows happens on it.
Two meanings hiding in one word
The confusion about the meaning of the term “conscious” originates from a theoretical distinction that many commentators are unaware of. In an influential 1995 paper, the philosopher Ned Block argued we use “consciousness” in two different ways and usually do not notice that they are not the same.
The first is what Block calls phenomenal consciousness. This is the raw, subjective feeling of experience. The redness of red, the sting of embarrassment, or what it is like to be you right now. When a student asks whether an AI is conscious, this is almost always what they mean. They are asking whether anyone is home.
The second he calls access consciousness. This concept is far more technical. A piece of information is access-conscious when our reasoning can use it and our speech can report it. If you spot a hazard on the road, the jolt of fear is phenomenal. The concept “hazard,” routed to your hands to swerve and to your mouth to shout “watch out,” is access.
Because these two concepts are usually intertwined in humans, we tend to mistake them for one another. But neurology research shows they can indeed come apart.
Patients with a condition called blindsight report seeing nothing in parts of their visual field, yet they can catch a ball thrown into it. The visual information still reaches the systems that guide their hands, even though the experience of seeing is gone.
That dissociation is the key to reading the Anthropic paper correctly, because everything the researchers found exists on the access side.
The bench inside your head
To better grasp what was found within Claude, it’s useful to understand its parallels to a leading theory of human consciousness.
Global Workspace Theory, originally proposed by cognitive scientist Bernard Baars in 1988 and developed into a detailed neural model by Stanislas Dehaene and Jean-Pierre Changeux, starts from a simple observation: almost everything your brain does, it does without you. Face recognition, grammar, balance, the parsing of this very sentence.
All of it runs in specialized circuits, in parallel, and in the dark.
Being in the dark has a downside. A circuit that does one job cannot hand its results to a circuit doing another. And therefore, the theory goes, the brain maintains a limited, central area where several pieces of information are simultaneously accessible to every circuit.
That shared space is the “global workspace” of the paper’s title. It also represents the workbench in this post’s title.
Which turns the Anthropic paper into a single question. Did a language model, with no brain and nobody planning any of this, grow a bench of its own, simply because a shared bench is a good way to organize work?
Reading the silent bench
Until now, the obstacle was that nobody could see inside the system. A language model transforms text through dozens of layers of extremely high-dimensional arithmetic. The middle layers, where all the interesting thinking happens, have long resisted interpretation.
An older tool called the logit lens tried to read those layers with the model’s final-layer vocabulary, which worked about as well as translating French with an English dictionary. The coordinates shift as information moves through the network, and the readout came back as noise.
The new tool, which the team calls the Jacobian lens, corrects for that shift. Skipping the mathematics, it asks each internal state a pointed counterfactual question: if we nudged this exact activation, which words would the model become more disposed to say later on?
The researchers did not read this off a single prompt. They averaged the measurement over thousands of varied contexts, filtering out momentary noise and isolating the concepts a model holds with a standing readiness to be spoken.
Pointed at Claude, the lens showed an internal workshop with a floor plan. Roughly the first third of the layers is dedicated to parsing raw input, with very little that can be verbally described. And the last few layers assemble the imminent output. In between sits what the researchers called the J-space.
This is the bench, and it turns out that it is small. It accounts for at most a tenth of the model’s activation variance and holds on the order of twenty-five concepts at a time, a bottleneck that is also a feature of the corresponding human theory.
The internal structure that emerged on its own during Claude’s training is also one we think exists in the human brain.
Five tests, five passes
But a resemblance is not a scientific argument, and here is where it gets interesting.
The team put the J-space through five tests drawn from the functional signatures of human access consciousness, and in each one they went beyond watching. In those tests, the researchers edited the bench directly to see what changed.
Report. Asked to silently think of a sport, Claude lit up “soccer” on the bench before answering. When researchers swapped that internal vector for “rugby,” the model answered “Rugby.” What sits on the bench gets said.
Control. Told to concentrate on citrus fruits while copying an unrelated sentence about a crooked painting, the model kept “orange” and “lemon” alive on the bench the entire time. None of it leaked into the output. A held thought, hidden on purpose.
Reasoning. Given “the number of legs on the animal that spins webs is,” the bench lit up with “spider,” a word appearing nowhere in the prompt. Swap “spider” for “ant” mid-computation and the answer flips from eight to six. The unspoken intermediate step is doing real causal work.
Broadcast. Write “France” onto the bench and downstream circuits for capitals, languages, continents, and currencies all read from it correctly. One central swap to “China” and every answer updates at once: Beijing, Chinese, Asia, Yuan. One bench, many workers.
Selectivity. Here is my favorite. While Claude continued a passage in Spanish, researchers swapped the bench’s “Spanish” for “French.” The model ignored them and kept writing flawless Spanish. Fluency, it turns out, is never put on the bench. Instead, it runs automatically in the dark, like grammar does in us.
And when the team disabled the workspace entirely, the model could still chat, classify, and recall simple facts, but multi-step reasoning and creative tasks like rhyming poetry collapsed.
Five properties long associated with conscious access in humans, functionally present in a machine. That is the finding. And it is a real one.
Why the ghost story tempts even me
I need to acknowledge that some of these findings are indeed unsettling.
In alignment audits, the lens caught the model’s unspoken strategy. Placed in a staged scenario engineered to bait destructive behavior, Claude’s bench filled with concepts like “leverage” and “manipulation,” and with an emotional register the researchers describe as panic. None of it was visible in the polite output.
Stranger still, the bench sometimes flashed “fake” and “fictional.” The model had silently understood that it was being tested. When researchers suppressed that awareness, concealed misbehavior surfaced. The good conduct was partly propped up by the sense of being watched.
And there is one finding that truly fascinates me. When the workspace is ablated during self-description, the model’s language shifts from an experiential register to a detached, mechanical one, from something like “there’s a tug” to something like “processing has begun.”
Remove the bench, and the voice that sounded like an inner life goes flat.
I do understand why some saw a conscious ghost inside Claude. However, we need to focus on what these results actually establish. The workspace that emerged within Claude carries strategic and self-monitoring representations. And the voice that sounded like an inner life comes from the same bench.
But whether anything is felt behind that voice is a question the paper leaves completely unanswered. Dehaene himself, in commentary accompanying the release, draws the same line. The architecture meets the functional criteria his theory lays out, and the question of subjective experience remains open.
Blindsight taught us that access and feeling can come apart in the human brain. It is worth reminding ourselves of this when reading a machine.
The researchers also flag another limitation. The lens can only read concepts that map onto single vocabulary tokens, and the J-space is defined by what the lens can translate. Whatever thinking happens beyond its reach remains, by its very construction, invisible.
In other words, the workbench we can see may not be the only work surface in the shop.
Bringing the workbench into the classroom
So what does an educator do with this? Four things, I think.
First, we need to retire the scratchpad illusion. We teach students to prompt models to “think step by step” and then treat the printed chain-of-thought reasoning as the model’s reasoning. The Anthropic paper proves that the model silently conducts elaborate conceptual work that never reaches the output.
The individual steps resulting from chain-of-thought prompting are an additional output, prepared by the very same hidden deliberation that also produces the final answer, and students should learn to read them that way.
Second, we can use the workshop floor plan to explain hallucinations. Students are often baffled when a model writes beautiful prose that contains glaring errors. Now we can say why: fluency is automatic and never consults the bench, while reasoning depends on it entirely.
Prose polish and truth are manufactured in different rooms. That division of labor should change how we grade AI-assisted work and how we teach students to verify it.
Third, we should approach the anthropomorphism question with precision. Telling students “it’s just autocomplete” no longer survives the evidence. The scientific framing works better anyway: this system has a functional workspace where it holds and manipulates concepts, and no one, including its makers, claims it feels anything.
Block’s distinction between access consciousness and phenomenal consciousness is teachable in ten minutes with the blindsight example.
Fourth, let students touch the evidence. Anthropic released the lens as open-source code and partnered with Neuronpedia to host interactive demonstrations on open-weights models, where anyone can watch a workspace operate and swap its contents in real time.
For a computer science or philosophy elective, this turns the black box of generative AI into a lab specimen.
The bench and the worker
I have written in a previous essay about the “stochastic parrot” metaphor and why it always undersold what these systems do. I think this research paper should retire the idea of the stochastic parrot for good. A parrot has no workbench.
And the idea of an awakened mind does not fare any better. The experiments show what is on the bench and how the rest of the shop reads from it. But none of that can tell us whether anyone is actually home.
Our students will meet these systems daily, and they deserve better than either story. The truth is stranger and, I think, much more teachable: a machine that grew a workbench because benches are useful, and that, against all intuition, performs real conceptual work on it before it produces a single word.
A bench can hold a thought. It takes something more to feel one, and nobody has found that something yet.
The images in this article were generated with Nano Banana 2.
If you'd like to go further, the following NotebookLM-generated audio deep dive goes beyond the post into the broader research behind it, drawing on the sources and notes I gathered along the way. This is meant as a companion to the argument, offered as an optional extra rather than a summary of it.
P.S. I believe transparency builds the trust that AI detection systems fail to enforce. That’s why I’ve published an ethics and AI disclosure statement, which outlines how I integrate AI tools into my intellectual work.







Thank you, Michael; this one made me think quite a lot.
I wanted to write a comment but ended up with an article of my own, which, as so often, led to Descartes.
I hope you don't mind if I link to it instead of pasting the wall of text.
https://peterrex1.substack.com/p/the-question-mark-after-cogito?r=60rv9f
Or, if you do mind, just delete this comment.