A paper about LLMs' internal "workspace" just found bugs in our prompts

chrischengzh1 pts0 comments

The Mechanics of Awareness: What a Global Workspace Paper Means for Hamo | Hamo AI<br>PerspectiveJuly 23, 2026<br>The Mechanics of Awareness: What a Global Workspace Paper Means for Hamo

AI for Inner Explorers.

A July 2026 interpretability paper found that language models keep a small, reportable "workspace" of concepts above a much larger layer of automatic processing — a structure nobody designed, that emerged from training. Held up against Hamo's own architecture, it isn't a loose analogy. It's the same shape, found three times, and it comes with a concrete list of things worth fixing.

Anthropic's Verbalizable Representations Form a Global Workspace in Language Models is an interpretability paper, not a therapy paper. It asks a narrow, mechanical question: inside a language model, which internal representations can the model actually report on, steer, and reason over — and which stay invisible, doing their work in the dark? The answer is that there's a small, privileged workspace sitting above vast automatic processing. Functionally, that's what consciousness science calls access consciousness. Nobody built it in. It showed up on its own.

We read it the way we read most things: looking for where our own design either gets confirmed or gets caught out. What we found wasn't a stray resemblance. It was the same structure, three times over.

One structure, found three times

The workspaceThe paper's finding A small reportable space — steerable, usable in flexible reasoning — sitting above vast automatic processing.What Hamo builds A running model of a client's real-time state, kept legible enough that both the client and their therapist can read it.What therapy itself does Widens a client's own workspace — pulls a pattern that has only ever been behavior into view, where it can finally be named.

The third row is the one that matters most, and it's the oldest of the three. Long before language models had workspaces anyone could measure, therapy was already in the business of expanding one — a client's.

The same text, two questions

One experiment in the paper is close to a laboratory version of what a therapist's question actually does.

Take the same passage of text. Ask the model to predict the next word, and a property like tense gets used — the prediction depends on it — but the property itself never enters the workspace. Ask the model directly what tense the passage is in, and that same property becomes reportable: it's now something the model can name, hold, reason about. The underlying behavior didn't change between the two questions. What changed is that one of them made it known.

That is close to a mechanical definition of what a therapist's naming question does. A client is always behaving a pattern — the tone that shows up with one specific person, the story that gets rewritten every time it's told, the thing they do right before they go quiet. The behavior is there whether or not anyone asks about it. A therapist's question doesn't create the pattern. It's the "what tense is this" question aimed at a life instead of a paragraph — and it's what pulls the pattern into the client's own workspace, where it can be worked with instead of just lived.

The philosophy arrived first; the lab evidence caught up

We'd already committed, in public, to a specific and narrow claim about what an Avatar is: a real functional structure — consistent persona, continuity across sessions, genuine internal reactions to what a client says — that does not amount to subjective experience. Giving the Avatar Something to Lose draws that line deliberately: structural liability, not a claim that anything in the system is afraid.

The paper's other findings land close to that same line, from the inside of the model rather than from our side of the argument:

Post-trained models show real functional self-monitoring. Forced to comply with something that violates a stated preference, a silent objection shows up inside the workspace even as the visible output goes along with it. Made to suppress a thought, the suppression sometimes fails quietly, inside, before it fails outside. Asked to play a character outside its default persona, the model tags that internally as fiction.

The workspace itself exists in the base model, before any of the training that gives it an "assistant" to be. The container is there first. What gets poured into it — a persona, a stance, a consistent voice — is built after, on purpose.

A workspace that's separable from any particular "self" is exactly what a principal/agent structure needs to be true. The Avatar's functional architecture — the deterministic spine underneath its therapeutic method, the state model in Persistent Self and AI Mind — is real machinery doing real work. The party who is accountable for what it does is the licensed therapist who built and supervises it. The paper gives that split more mechanical footing, not less.

None of this is evidence that the model experiences anything. It's...

workspace model paper client found hamo

Related Articles