A test for AI consciousness
A test for AI consciousness
Aug 16, 2026
Whether AI is, or ever will be, conscious is hotly contested.<br>Simply asking the models can never resolve the question definitively.<br>For whatever reason, they may pretend to be conscious when they actually aren't,<br>or pretend not to be when they actually are.<br>Claude, for that matter, is explicitly instructed to dodge the question.<br>Some people believe the answer is fundamentally unknowable.<br>But I want to give it a shot and propose a test for AI consciousness1.
A philosophical zombie is a creature that seems to act and behave like a conscious being but lacks conscious experience2.<br>AI models may be philosophical zombies.
Philosophical zombie model
According to epiphenomenalism, consciousness is purely passive.<br>Like a person watching a movie.<br>All sensory input from the outside world and all internal thoughts and feelings are playing out on screen.<br>The viewer is so immersed that they have the illusion of control.<br>In this picture, no information flows back from consciousness into the "physical world".<br>Hence, from the outside there is no way to distinguish a conscious being from a philosophical zombie.
Epiphenomenalist model
But this picture can't be true.<br>At least in one way consciousness has a measurable effect on the world: we talk about it.<br>Talking is a physical process orchestrated by the excitement of neurons in the brain.<br>But what is initiating this process?<br>Where is the initial spark coming from?<br>It might come from a "physical stimulus": a book, a conversation, a memory of a conversation.<br>That's why philosophical zombies can surely engage in a dialog about consciousness.<br>But absent any physical stimulus, the spark must come from consciousness itself.<br>At least sometimes consciousness must feed directly into the excitement of neurons and compel us to talk about it.<br>A lone philosophical zombie in a population of conscious beings might fly under the radar,<br>but a population purely of philosophical zombies would look different than ours.<br>They could never have initiated the discussion about consciousness and spilled all that ink.
The question is how to eliminate the confounding factor of physical stimuli.<br>We could train a model on a dataset where any reference to consciousness is thoroughly removed.<br>Then explore whether the model can bootstrap a discussion about consciousness.
How practical this test is, is another question.<br>Scrubbing the dataset is non-trivial.<br>The topic is discussed all over the place in philosophy, religion, psychology, medicine, fiction, etc.<br>If that's done successfully, the model won't have the very language to talk about consciousness.<br>It would have to invent its own.<br>Also, the test can arguably only confirm existence of consciousness but not rule it out.<br>If the model never declares to be conscious, the reason could be unrelated.
Nevertheless, I think this is a decent case that the question is not fundamentally unresolvable.
Footnotes
Specifically as in the hard problem of consciousness.
Not the usual definition but I need a term that: also covers AI models and is not by definition indistinguishable from the conscious counterpart.