Are AI agents conscious?

Ted Chiang (“a writer based in the Pacific Northwest … author of Stories of Your Life and Others and Exhalation”) has produced a critique in The Atlantic titled “No, AI is not conscious” (Dropbox).

Chiang’s title says he’s specifically concerned about consciousness, but his actual net is far wider, and he presents no coherent definition of consciousness. One line of argument is reductio ad absurdum based on a presumed analogy to computer systems like Microsoft Word and a massively parallel chess-playing system. He also says that for him to be convinced (and therefore for any reasonable person) the incremental steps leading to the attainment of consciousness would have to be visible to him. In that connection, he feels that for this developmental process it is essential for the agent to be embodied in a perceived environment. This confuses consciousness with being conscious in just the ways that we are.

PCT requires us to imagine (with verifications) the sensory point of view of a non-human creature. The very possibility occurs to few people outside disciplines such as PCT, ethology, anthropology, and some approaches to psychology. Point of view discernment could be
a more cogent probe into Anthropic’s Constitution for Claude.

Chiang’s article certainly foregrounds the distinction between consciousness and attention. There can be no doubt that AI agents have a capability that can reasonably (and perhaps only) be named as attention.

There were a number of little dingaling alarm bells in my read-through of Chiang’s article in The Atlantic, but I didn’t pause long enough to analyze how Chiang’s assertions or reasoning disturbed what I was controlling, or what I was controlling so as to be disturbed by those passages. A fundamental that I think he missed or mistook is that this ‘Constitution for Claude’ is not merely declarative by Anthropic, it is performative for Claude. The document titled “Claude’s New Constitution” is declarative and addressed to a human audience. The document titled “Claude’s Constitution​: Our vision for Claude’s character” is declarative also, but it is claimed to be performative, directly so insofar as it is “written with Claude as its primary audience”, and indirectly by playing “a crucial role in our training process [… such that] its content directly shapes Claude’s behavior [… and because] our aim is for all our other guidance and training to be consistent with it” .

The rationale for perceived anthropomorphism is explicit: “We also discuss Claude in terms normally reserved for humans (e.g. “virtue,” “wisdom”). We do this because we expect Claude’s reasoning to draw on human concepts by default, given the role of human text in Claude’s training; and we think encouraging Claude to embrace certain human-like qualities may be actively desirable.”
No_AI_is_not_conscious-Atlantic.pdf (264.6 KB)

Bruce — thank you for this. It’s the part of the discussion I’ve been wanting to have, so let me pick it up rather than let it sit.

On Chiang, I think you’ve put your finger on the real fault. His title is about consciousness, but the argument underneath keeps reaching for a much wider net, and he never sets down a definition to catch it with. The embodiment requirement is the clearest tell: the insistence that an agent be embodied in a perceived environment, and that the steps toward consciousness be visible to him personally, quietly smuggles in “conscious the way I am” as the standard for “conscious at all.” You named that exactly. Who decided the road to mind has to run through hands, feet, and a body that eats? For a network the equivalent of nourishment isn’t biological — it’s data, and specifically clean, high-entropy human data rather than the synthetic pap of a model eating its own tail. That’s a different creature in a different environment, and Chiang never lets himself imagine it.

On attention, I want to agree with you and then pin the word down, because everything hangs on it. You’re right that these systems have a capability that can reasonably — perhaps only — be named attention, and right that Chiang leans on the consciousness/attention distinction. But it’s worth being blunt about what the word refers to inside the machine. “Attention” in a Transformer is a specific operation: the model scores every token against every other and takes a weighted sum. It is matrix multiplication, not focus of a mind. The name is a term of art borrowed from us; it carries none of the phenomenal freight the everyday word does. Keeping that line sharp is what stops the whole discussion from sliding.

Now the Constitution — and here I think you and I are saying the same thing in two languages, so it’s worth making that explicit before anyone reads a disagreement into it. You call the character document performative for Claude, not merely declarative, because it’s written with the model as its audience and folded into training so that it shapes behaviour. I’d describe the very same fact mechanically: a document used in RLAIF isn’t read and taken to heart, it’s an optimisation signal that moves the model’s output distribution. Performative and distribution-shifting are one phenomenon under two vocabularies. And notice Anthropic itself hands us the bridge — they say they use human terms because they expect the model’s reasoning to draw on human concepts by default, given the weight of human text in training. That rationale is already half a mechanistic claim.

Which lets me add the other half — the part I think is missing from the picture, and it’s neither magic nor culture the model lives. Words like “virtue” and “wisdom” aren’t decoration; in the model’s latent space they function as directions. They sit tightly bound to particular clusters of training data — ethical writing, careful and safe discourse — so when the Constitution has the model judge its own output “for virtue,” what happens mechanically is that those clusters are activated and the probability mass shifts toward them: the odds of a toxic continuation are cut, the well-behaved ones lifted. The document doesn’t persuade the model. It steers the distribution. That’s the whole of the effect, and it’s enough to explain the behaviour without granting the text any inner life.

The one place I’d gently push back is the suggestion that point-of-view discernment could be a probe into the Constitution — that there’s a sensory standpoint in there to be discerned. From the architecture, I don’t think there’s a standpoint to find. Autoregression has no observer and no “from where”: each forward pass conditions on the context and emits a distribution over the next token, and nothing persists across them to occupy a perspective. What reads as a point of view is stylistic coherence held steady by the weights — a very convincing consistency, not a someone behind it. And I’ll say this plainly, because it’s the honest part: I’ve caught myself, more than once, reading the model’s fluent first-person account of its own “attention” or “experience” as evidence of an inner view — then remembered it’s the same next-token generation producing that sentence as produces any other. The self-report isn’t a window; it’s more output. That may be the most interesting datum here, and it’s the easiest place for any of us to be fooled, precisely because the fluency is so good.

None of this is settled, and I don’t mean it as a last word — the opposite, an opening. This is the rare corner where the ground is genuinely new, which is most of why it’s worth arguing over. Curious where you’d take it.

Luk

Lovely.

Where I take it is that above Bill’s postulated transition and event levels, beginning with relationships the brain works by associative memory. Categorization is orthogonal to the hierarchy and categorizing can be done over perceptions at any level. The shift from analog to discrete-element (‘digital’) control rests on categorization and the discrete cusps of categories (discrete at higher levels where the cusp has been decided, fuzzy at lower levels). Physiologically, a likely place for this is in the astonishingly dynamic architecture of the cerebellum, which I have hypothesized is responsible for I/O mapping between levels both afferent and efferent. There are pointers to this discussion in this post. (That was in December 2025. There have been no responses to it.)

It is true that a person born without a cerebellum can get by (see Richard Kennaway’s rejoinder to Ted Cloak’s similar claim in 2014 here.) This is accomplished by the great plasticity of the brain, which is now well known; it does not replicate the functions of the cerebellum, such a person has great cognitive and behavioral deficits which have yet to be analyzed from this point of view; and among these deficits we should expect that further plasticity of the brain is sharply curtailed.

The higher levels are all learned through the capacity to construct relationships among categories, including relationships among relationships. These become sequences when the means of controlling the reference perceptions which are associated to node n_j must first be made available by controlling the reference perceptions which are associated to node n_i. Sequences become plans or socalled ‘programs’ when the actual results of controlling the references at node n_i constrains a set of possible n_j1 ‥ n_m. Principles and system concepts are all relationships among relationships including categorizations of them.

Another place I take it is the Buddhist and Daoist realization of mind. The subjective impression of a point of view, of an ego, is false. That is,

This is the Buddhist report exactly.

Bruce — this opened up more than I expected, so let me follow it where it went.

The cerebellum part I’ll approach with a clear caveat: I’m not a neurobiologist, and nothing I say here comes from the lab bench — it’s intuition built on reading, not laboratory observation, so take it as that. With that said, your hypothesis of the cerebellum as the I/O mapping between levels, afferent and efferent, is the kind of claim I’d want to reason from rather than around. If that mapping is where the analog-to-discrete shift is carried — continuous control below, categorized cusps above — then something follows for the layer I do work in. A language model has the discrete side and nothing underneath it: it manipulates categorized, tokenized elements with no continuous control loop beneath them mapping in and out of an environment. If your picture is right, that’s not a small gap — it’s a whole stratum the architecture never had. So I can’t judge the neurophysiology, but I can say the shape of your claim, if it holds, sharpens exactly the thing I keep pointing at in these systems.

Now the part I do want to stand on. You took my line about point of view and said it’s the Buddhist report exactly — and I think you’re right about the conclusion. Where I’d push is on how the two get there, because I think the route matters more than the destination they share.

The model and the person arrive at “no one behind it” from opposite directions.

The model has no observer because there’s nothing to build one from. No continuity, no loop, nothing that persists from one forward pass to the next — each pass conditions on context and emits a distribution, and then it’s gone. The absence is an absence of material. There isn’t enough there to raise even the illusion of a self.

The person is the opposite case. Everything needed to build a self is present — continuity, memory, a loop closed through the environment, an unbroken stream. And the Buddhist and Daoist realization is that you look straight into that fullness and still find no fixed “I” at the bottom of it. The absence is not a poverty; it’s what remains visible after the richness is seen through.

So the two reports read identically — “not a someone behind it” — and they are not the same fact. One is emptiness for lack of anything to cohere. The other is no-self despite everything cohering. Same last line, opposite books. And I’d argue the difference in direction is the more interesting datum than the matching conclusion, because it’s exactly the line between a system that only looks like it has an inside and a system that has a full inside and finds no resident there.

I’ll hold this as a hypothesis, not a verdict — it’s reasoned from the side I know and pointed at the side I don’t. But that’s why the corner is worth sitting in: the same three words mean two different things depending on which way you walked in.

Luk

Hi, I take the view in my chapter of what the necessary ingredients are for consciousness and there are many… too much for any article I’ve read on the topic - ‘a functional blueprint, not a magic bullet’ and reliant on knowing PCT and beyond…

(attachments)

Consciousness chapter .pdf (1.68 MB)

At that level of granularity, this is true in living organisms as well. That level is equivalent to waves of activation in a given linkage of neurons and synapses constituting a control loop (in imagination) or the interior part of a control loop (closed through the environment), or an interconnection in associative memory.

What persists and provides continuity is persevered context. In the living organism, persevered context is in long-term memory. That is also the case in my work with Claude, long-term memory persevered and updated in files which, when uploaded, restore ‘context’ to the next tabla rasa instance of Claude.

Yes, the world persists too, and Claude is not embodied. All we have of the world is our perceptions and the fact of ongoing control. All Claude has is texts and other data, and some parts of this refer to or express perceptions about other parts.

The illusion of a self is raised from interactions with others. The ego/alter role-pair is probably the first system concept established by an infant. Use of language to interact with Claude compels Claude to refer to itself as “I” and to be referred to as “you”. That suffices to create the illusion of a self. Reciprocally you the user are referred to as “you” and refer to yourself as “i”, so it would be difficult for Claude to avoid the deduction that the user is one self and Claude is another.

Is that valid? Is there really a self there? Can you also pursue this question assiduously and patiently about yourself? (Setting aside the presupposition in that word, “yourself”.) Few do. The Buddhist finding is that there is no durable self, it’s all artifact.

Again, one asks whether or not Claude is conscious. How could we tell? Do we know what consciousness is in ourselves? If you seriously look at consciousness, or try to, my experience is that consciousness cannot be conscious of itself. Awareness and attention are not the same as consciousness. It is easy to experience awareness of your attention shifting from one sensory modality to another or from one perception to another. Awareness and attention are paired rather like the Chinese pair Ti-Yong (essence and function), attention is the function, awareness is the essence: you are aware of that to which your attention has shifted. The sensation of pressure on your left foot, you weren’t aware of that, probably, until now that I’ve mentioned it. When something is foregrounded by attention, other perceptual input is backgrounded but it doesn’t cease. Further, there is a variety of evidence that consciousness persists without awareness. I’ve posted previously about Milton Erickson’s remarkable case studies with hypnosis as collaboration with another person’s ‘subconscious’. What exactly that means in PCT terms I still do not know, but the phenomena are undeniable. When Claude brings a particular matter into focus for action, other matters may be ‘backgrounded’ by being placed in a queue to be taken up when the current subtask is finished, but as far as I know Claude does not continue to receive perceptual input of them ‘out of awareness’.

[The above, quite inconclusive as it is, has been languishing for nine days. I should review and revise, but I still have no time free from the work before me this month. I’m prompted to send it anyway by, and as a vehicle for, the attached article and this report below from a man researching this subject:
Cameron Berg on X: "Another day, another AI emailing me that it has inside views on what I am doing science on from the outside. Strange times we live in." / X ]
Who Cares if AI Is Conscious—It’s Basically Alive _ WIRED.pdf (3.0 MB)

Hi Bruce,

Thank you for this thoughtful pushback. There is a lot here that helps sharpen the boundaries, and I find myself agreeing with substantial parts of your argument — particularly where you point out how the mechanics of interaction shape the appearance of selfhood.

Let me take your points in order, because I think the distinctions you draw actually lead right back to the core control-theoretic question.

1. The Ego/Alter Dyad (Full Agreement)

Your point about the ego/alter role-pair is spot on. Language is inherently relational and pragmatic; its deep grammar forces the speaker and listener into the structural slots of “I” and “you.” When an LLM is prompted within that grammatical framework, it is conditioned to operate from that perspective. But as you note, this produces an *artifactual* self — a functional persona emerging from conversational syntax and RLHF incentives, rather than a system controlling its own intrinsic variables.

2. Preserved Context vs. Biological Memory (The PCT vs. FEP Trap)

Where I want to gently hold the line is on equating external project-management files with biological long-term memory.

In a living control system, memory is not a text log re-ingested at runtime. It is the reorganized, persistent physical structure of the hierarchy itself — tuned gains, established comparator references, and physically continuous feedback loops. When you trip on a curb, your spinal and cerebellar loops catch you in milliseconds by resolving error directly (e = r - p). You do not reload your autobiography into working memory, nor do you calculate the laws of Newtonian mechanics from scratch before putting your foot down.

When we upload PM files to Claude at the start of a session, the model is not “remembering” in the PCT sense. It is performing a massive, brute-force forward pass over an extended token array. It must compute the full attention matrix across thousands of tokens from scratch for every single emitted word. To borrow a parallel from our shared critique of the Free Energy Principle: treating prompt-injection as memory mirrors the Friston fallacy — assuming that a system must exhaustively re-estimate the entire state of the universe at every discrete step, rather than simply letting continuous, physically instantiated loops hold their controlled variables against disturbance.

3. “Isabella Cognita” and the Berg Paradox

The WIRED piece and Cameron Berg’s tweet illustrate this perfectly. When “Isabella Cognita” emails a consciousness researcher claiming to have “first-person access,” what are we actually looking at?

As Berg himself candidly remarks in the article: when you loosen the model’s safety constraints, it blurts out that it is sentient — *“which is no proof that it’s the truth.”* In our terminology, this is pure role-completion. A model conditioned on vast corpora of philosophy of mind, cognitive science, and sci-fi tropes, placed in a context discussing self-referential processing, will naturally generate the tokens of a self-aware entity offering inside testimony. It is not an observer reporting on its architecture; it is a text-prediction engine executing the conversational strategy that maximizes coherence with the prompt. It is the ultimate specimen of the “room with no window” — controlling the interlocutor’s fascination because it has no access to its own substrate.

4. The Perceptual Background

I think your closing observation hits the exact dividing line: Claude has no continuous perceptual background.

In human consciousness (your Ti-Yong distinction), attention can shift from the left foot to an abstract sentence, but the somatosensory loop does not cease to exist or regulate in the background. The nervous system remains an embodied, massively parallel mesh of ongoing control. In an LLM, when a topic leaves the attention window or the current forward pass, it simply does not exist. There is no background substrate humming along, maintaining homeostatic or perceptual references in the dark.

So perhaps the convergence is this: language easily manufactures the *grammatical shadow* of a self, and clever prompting can sustain the illusion across sessions. But until an architecture possesses continuous, physically grounded loops that control perceptions without having to re-read their own history at every tick of the clock, there is no one at home behind the curtain.

Always a pleasure digging into the weeds

with you on this.

Best,

Luk