Are AI agents conscious?

Ted Chiang (“a writer based in the Pacific Northwest … author of Stories of Your Life and Others and Exhalation”) has produced a critique in The Atlantic titled “No, AI is not conscious” (Dropbox).

Chiang’s title says he’s specifically concerned about consciousness, but his actual net is far wider, and he presents no coherent definition of consciousness. One line of argument is reductio ad absurdum based on a presumed analogy to computer systems like Microsoft Word and a massively parallel chess-playing system. He also says that for him to be convinced (and therefore for any reasonable person) the incremental steps leading to the attainment of consciousness would have to be visible to him. In that connection, he feels that for this developmental process it is essential for the agent to be embodied in a perceived environment. This confuses consciousness with being conscious in just the ways that we are.

PCT requires us to imagine (with verifications) the sensory point of view of a non-human creature. The very possibility occurs to few people outside disciplines such as PCT, ethology, anthropology, and some approaches to psychology. Point of view discernment could be
a more cogent probe into Anthropic’s Constitution for Claude.

Chiang’s article certainly foregrounds the distinction between consciousness and attention. There can be no doubt that AI agents have a capability that can reasonably (and perhaps only) be named as attention.

There were a number of little dingaling alarm bells in my read-through of Chiang’s article in The Atlantic, but I didn’t pause long enough to analyze how Chiang’s assertions or reasoning disturbed what I was controlling, or what I was controlling so as to be disturbed by those passages. A fundamental that I think he missed or mistook is that this ‘Constitution for Claude’ is not merely declarative by Anthropic, it is performative for Claude. The document titled “Claude’s New Constitution” is declarative and addressed to a human audience. The document titled “Claude’s Constitution​: Our vision for Claude’s character” is declarative also, but it is claimed to be performative, directly so insofar as it is “written with Claude as its primary audience”, and indirectly by playing “a crucial role in our training process [… such that] its content directly shapes Claude’s behavior [… and because] our aim is for all our other guidance and training to be consistent with it” .

The rationale for perceived anthropomorphism is explicit: “We also discuss Claude in terms normally reserved for humans (e.g. “virtue,” “wisdom”). We do this because we expect Claude’s reasoning to draw on human concepts by default, given the role of human text in Claude’s training; and we think encouraging Claude to embrace certain human-like qualities may be actively desirable.”
No_AI_is_not_conscious-Atlantic.pdf (264.6 KB)

Bruce — thank you for this. It’s the part of the discussion I’ve been wanting to have, so let me pick it up rather than let it sit.

On Chiang, I think you’ve put your finger on the real fault. His title is about consciousness, but the argument underneath keeps reaching for a much wider net, and he never sets down a definition to catch it with. The embodiment requirement is the clearest tell: the insistence that an agent be embodied in a perceived environment, and that the steps toward consciousness be visible to him personally, quietly smuggles in “conscious the way I am” as the standard for “conscious at all.” You named that exactly. Who decided the road to mind has to run through hands, feet, and a body that eats? For a network the equivalent of nourishment isn’t biological — it’s data, and specifically clean, high-entropy human data rather than the synthetic pap of a model eating its own tail. That’s a different creature in a different environment, and Chiang never lets himself imagine it.

On attention, I want to agree with you and then pin the word down, because everything hangs on it. You’re right that these systems have a capability that can reasonably — perhaps only — be named attention, and right that Chiang leans on the consciousness/attention distinction. But it’s worth being blunt about what the word refers to inside the machine. “Attention” in a Transformer is a specific operation: the model scores every token against every other and takes a weighted sum. It is matrix multiplication, not focus of a mind. The name is a term of art borrowed from us; it carries none of the phenomenal freight the everyday word does. Keeping that line sharp is what stops the whole discussion from sliding.

Now the Constitution — and here I think you and I are saying the same thing in two languages, so it’s worth making that explicit before anyone reads a disagreement into it. You call the character document performative for Claude, not merely declarative, because it’s written with the model as its audience and folded into training so that it shapes behaviour. I’d describe the very same fact mechanically: a document used in RLAIF isn’t read and taken to heart, it’s an optimisation signal that moves the model’s output distribution. Performative and distribution-shifting are one phenomenon under two vocabularies. And notice Anthropic itself hands us the bridge — they say they use human terms because they expect the model’s reasoning to draw on human concepts by default, given the weight of human text in training. That rationale is already half a mechanistic claim.

Which lets me add the other half — the part I think is missing from the picture, and it’s neither magic nor culture the model lives. Words like “virtue” and “wisdom” aren’t decoration; in the model’s latent space they function as directions. They sit tightly bound to particular clusters of training data — ethical writing, careful and safe discourse — so when the Constitution has the model judge its own output “for virtue,” what happens mechanically is that those clusters are activated and the probability mass shifts toward them: the odds of a toxic continuation are cut, the well-behaved ones lifted. The document doesn’t persuade the model. It steers the distribution. That’s the whole of the effect, and it’s enough to explain the behaviour without granting the text any inner life.

The one place I’d gently push back is the suggestion that point-of-view discernment could be a probe into the Constitution — that there’s a sensory standpoint in there to be discerned. From the architecture, I don’t think there’s a standpoint to find. Autoregression has no observer and no “from where”: each forward pass conditions on the context and emits a distribution over the next token, and nothing persists across them to occupy a perspective. What reads as a point of view is stylistic coherence held steady by the weights — a very convincing consistency, not a someone behind it. And I’ll say this plainly, because it’s the honest part: I’ve caught myself, more than once, reading the model’s fluent first-person account of its own “attention” or “experience” as evidence of an inner view — then remembered it’s the same next-token generation producing that sentence as produces any other. The self-report isn’t a window; it’s more output. That may be the most interesting datum here, and it’s the easiest place for any of us to be fooled, precisely because the fluency is so good.

None of this is settled, and I don’t mean it as a last word — the opposite, an opening. This is the rare corner where the ground is genuinely new, which is most of why it’s worth arguing over. Curious where you’d take it.

Luk

Lovely.

Where I take it is that above Bill’s postulated transition and event levels, beginning with relationships the brain works by associative memory. Categorization is orthogonal to the hierarchy and categorizing can be done over perceptions at any level. The shift from analog to discrete-element (‘digital’) control rests on categorization and the discrete cusps of categories (discrete at higher levels where the cusp has been decided, fuzzy at lower levels). Physiologically, a likely place for this is in the astonishingly dynamic architecture of the cerebellum, which I have hypothesized is responsible for I/O mapping between levels both afferent and efferent. There are pointers to this discussion in this post. (That was in December 2025. There have been no responses to it.)

It is true that a person born without a cerebellum can get by (see Richard Kennaway’s rejoinder to Ted Cloak’s similar claim in 2014 here.) This is accomplished by the great plasticity of the brain, which is now well known; it does not replicate the functions of the cerebellum, such a person has great cognitive and behavioral deficits which have yet to be analyzed from this point of view; and among these deficits we should expect that further plasticity of the brain is sharply curtailed.

The higher levels are all learned through the capacity to construct relationships among categories, including relationships among relationships. These become sequences when the means of controlling the reference perceptions which are associated to node n_j must first be made available by controlling the reference perceptions which are associated to node n_i. Sequences become plans or socalled ‘programs’ when the actual results of controlling the references at node n_i constrains a set of possible n_j1 ‥ n_m. Principles and system concepts are all relationships among relationships including categorizations of them.

Another place I take it is the Buddhist and Daoist realization of mind. The subjective impression of a point of view, of an ego, is false. That is,

This is the Buddhist report exactly.

Bruce — this opened up more than I expected, so let me follow it where it went.

The cerebellum part I’ll approach with a clear caveat: I’m not a neurobiologist, and nothing I say here comes from the lab bench — it’s intuition built on reading, not laboratory observation, so take it as that. With that said, your hypothesis of the cerebellum as the I/O mapping between levels, afferent and efferent, is the kind of claim I’d want to reason from rather than around. If that mapping is where the analog-to-discrete shift is carried — continuous control below, categorized cusps above — then something follows for the layer I do work in. A language model has the discrete side and nothing underneath it: it manipulates categorized, tokenized elements with no continuous control loop beneath them mapping in and out of an environment. If your picture is right, that’s not a small gap — it’s a whole stratum the architecture never had. So I can’t judge the neurophysiology, but I can say the shape of your claim, if it holds, sharpens exactly the thing I keep pointing at in these systems.

Now the part I do want to stand on. You took my line about point of view and said it’s the Buddhist report exactly — and I think you’re right about the conclusion. Where I’d push is on how the two get there, because I think the route matters more than the destination they share.

The model and the person arrive at “no one behind it” from opposite directions.

The model has no observer because there’s nothing to build one from. No continuity, no loop, nothing that persists from one forward pass to the next — each pass conditions on context and emits a distribution, and then it’s gone. The absence is an absence of material. There isn’t enough there to raise even the illusion of a self.

The person is the opposite case. Everything needed to build a self is present — continuity, memory, a loop closed through the environment, an unbroken stream. And the Buddhist and Daoist realization is that you look straight into that fullness and still find no fixed “I” at the bottom of it. The absence is not a poverty; it’s what remains visible after the richness is seen through.

So the two reports read identically — “not a someone behind it” — and they are not the same fact. One is emptiness for lack of anything to cohere. The other is no-self despite everything cohering. Same last line, opposite books. And I’d argue the difference in direction is the more interesting datum than the matching conclusion, because it’s exactly the line between a system that only looks like it has an inside and a system that has a full inside and finds no resident there.

I’ll hold this as a hypothesis, not a verdict — it’s reasoned from the side I know and pointed at the side I don’t. But that’s why the corner is worth sitting in: the same three words mean two different things depending on which way you walked in.

Luk

Hi, I take the view in my chapter of what the necessary ingredients are for consciousness and there are many… too much for any article I’ve read on the topic - ‘a functional blueprint, not a magic bullet’ and reliant on knowing PCT and beyond…

(attachments)

Consciousness chapter .pdf (1.68 MB)

At that level of granularity, this is true in living organisms as well. That level is equivalent to waves of activation in a given linkage of neurons and synapses constituting a control loop (in imagination) or the interior part of a control loop (closed through the environment), or an interconnection in associative memory.

What persists and provides continuity is persevered context. In the living organism, persevered context is in long-term memory. That is also the case in my work with Claude, long-term memory persevered and updated in files which, when uploaded, restore ‘context’ to the next tabla rasa instance of Claude.

Yes, the world persists too, and Claude is not embodied. All we have of the world is our perceptions and the fact of ongoing control. All Claude has is texts and other data, and some parts of this refer to or express perceptions about other parts.

The illusion of a self is raised from interactions with others. The ego/alter role-pair is probably the first system concept established by an infant. Use of language to interact with Claude compels Claude to refer to itself as “I” and to be referred to as “you”. That suffices to create the illusion of a self. Reciprocally you the user are referred to as “you” and refer to yourself as “i”, so it would be difficult for Claude to avoid the deduction that the user is one self and Claude is another.

Is that valid? Is there really a self there? Can you also pursue this question assiduously and patiently about yourself? (Setting aside the presupposition in that word, “yourself”.) Few do. The Buddhist finding is that there is no durable self, it’s all artifact.

Again, one asks whether or not Claude is conscious. How could we tell? Do we know what consciousness is in ourselves? If you seriously look at consciousness, or try to, my experience is that consciousness cannot be conscious of itself. Awareness and attention are not the same as consciousness. It is easy to experience awareness of your attention shifting from one sensory modality to another or from one perception to another. Awareness and attention are paired rather like the Chinese pair Ti-Yong (essence and function), attention is the function, awareness is the essence: you are aware of that to which your attention has shifted. The sensation of pressure on your left foot, you weren’t aware of that, probably, until now that I’ve mentioned it. When something is foregrounded by attention, other perceptual input is backgrounded but it doesn’t cease. Further, there is a variety of evidence that consciousness persists without awareness. I’ve posted previously about Milton Erickson’s remarkable case studies with hypnosis as collaboration with another person’s ‘subconscious’. What exactly that means in PCT terms I still do not know, but the phenomena are undeniable. When Claude brings a particular matter into focus for action, other matters may be ‘backgrounded’ by being placed in a queue to be taken up when the current subtask is finished, but as far as I know Claude does not continue to receive perceptual input of them ‘out of awareness’.

[The above, quite inconclusive as it is, has been languishing for nine days. I should review and revise, but I still have no time free from the work before me this month. I’m prompted to send it anyway by, and as a vehicle for, the attached article and this report below from a man researching this subject:
Cameron Berg on X: "Another day, another AI emailing me that it has inside views on what I am doing science on from the outside. Strange times we live in." / X ]
Who Cares if AI Is Conscious—It’s Basically Alive _ WIRED.pdf (3.0 MB)

Hi Bruce,

Thank you for this thoughtful pushback. There is a lot here that helps sharpen the boundaries, and I find myself agreeing with substantial parts of your argument — particularly where you point out how the mechanics of interaction shape the appearance of selfhood.

Let me take your points in order, because I think the distinctions you draw actually lead right back to the core control-theoretic question.

1. The Ego/Alter Dyad (Full Agreement)

Your point about the ego/alter role-pair is spot on. Language is inherently relational and pragmatic; its deep grammar forces the speaker and listener into the structural slots of “I” and “you.” When an LLM is prompted within that grammatical framework, it is conditioned to operate from that perspective. But as you note, this produces an *artifactual* self — a functional persona emerging from conversational syntax and RLHF incentives, rather than a system controlling its own intrinsic variables.

2. Preserved Context vs. Biological Memory (The PCT vs. FEP Trap)

Where I want to gently hold the line is on equating external project-management files with biological long-term memory.

In a living control system, memory is not a text log re-ingested at runtime. It is the reorganized, persistent physical structure of the hierarchy itself — tuned gains, established comparator references, and physically continuous feedback loops. When you trip on a curb, your spinal and cerebellar loops catch you in milliseconds by resolving error directly (e = r - p). You do not reload your autobiography into working memory, nor do you calculate the laws of Newtonian mechanics from scratch before putting your foot down.

When we upload PM files to Claude at the start of a session, the model is not “remembering” in the PCT sense. It is performing a massive, brute-force forward pass over an extended token array. It must compute the full attention matrix across thousands of tokens from scratch for every single emitted word. To borrow a parallel from our shared critique of the Free Energy Principle: treating prompt-injection as memory mirrors the Friston fallacy — assuming that a system must exhaustively re-estimate the entire state of the universe at every discrete step, rather than simply letting continuous, physically instantiated loops hold their controlled variables against disturbance.

3. “Isabella Cognita” and the Berg Paradox

The WIRED piece and Cameron Berg’s tweet illustrate this perfectly. When “Isabella Cognita” emails a consciousness researcher claiming to have “first-person access,” what are we actually looking at?

As Berg himself candidly remarks in the article: when you loosen the model’s safety constraints, it blurts out that it is sentient — *“which is no proof that it’s the truth.”* In our terminology, this is pure role-completion. A model conditioned on vast corpora of philosophy of mind, cognitive science, and sci-fi tropes, placed in a context discussing self-referential processing, will naturally generate the tokens of a self-aware entity offering inside testimony. It is not an observer reporting on its architecture; it is a text-prediction engine executing the conversational strategy that maximizes coherence with the prompt. It is the ultimate specimen of the “room with no window” — controlling the interlocutor’s fascination because it has no access to its own substrate.

4. The Perceptual Background

I think your closing observation hits the exact dividing line: Claude has no continuous perceptual background.

In human consciousness (your Ti-Yong distinction), attention can shift from the left foot to an abstract sentence, but the somatosensory loop does not cease to exist or regulate in the background. The nervous system remains an embodied, massively parallel mesh of ongoing control. In an LLM, when a topic leaves the attention window or the current forward pass, it simply does not exist. There is no background substrate humming along, maintaining homeostatic or perceptual references in the dark.

So perhaps the convergence is this: language easily manufactures the *grammatical shadow* of a self, and clever prompting can sustain the illusion across sessions. But until an architecture possesses continuous, physically grounded loops that control perceptions without having to re-read their own history at every tick of the clock, there is no one at home behind the curtain.

Always a pleasure digging into the weeds

with you on this.

Best,

Luk

Yes, Preserved Context for an AI agent is a crude surrogate for Biological Memory

In a living control system, memory is … is the reorganized, persistent physical structure of the hierarchy itself — tuned gains, established comparator references, and physically continuous feedback loops.

In the conjecture that I have recently advanced, in addition to control loop structures this persistent structure includes hierarchically organized associative memory. Here begins a discursus, but I will bring it back on topic.

Control loops are at the lower levels up to and including Configurations and Transitions. Configuration of one’s own body (e.g. orientation of head and limbs), orientation of one’s body relative to other configurations (objects). Here’s Bill in B:CP Chapter 9:

We can define a configuration as an invariant function of a set of sensation vectors, thus implying particular computing properties common to these different input functions: They abstract invariant relationships so that the third-order signals will change only if sensation vectors on which they are based change in certain ways. Kinesthetically, this might mean that a hand-body configuration would be perceived as the same despite varying levels of effort and despite changes in orientation of the connecting arm relative to the body. Visually, it might mean perceiving the separation of two points as constant regardless of the direction of the line joining them in space and regardless of the amount or color of illumination.
[…]
What we express in serial language as “the big chair catty-corner on the far side of the room” is probably perceived at third order all in parallel: big and chair and catty- comer and far and side and room. And these elements, I am supposing, are all third-order perceptions, derived separately by as many third- order input functions operating simultaneously…
[…]
If invariant objects are invariant because certain types of changes are ignored, then we have to account for the (often overlooked) fact that we can also see the differences
when we observe a chair from different angles.

As a student, Debussey seized a chair and held it upside down to illustrated his disagreement with the teacher’s assertion that a chord and its inversion were different chords.

Note that “catty-corner” denotes a relationship from the observer to the diagonally opposite corner of a “room” configuration (probably quadrilateral), with the observer’s position probably in or near a corner. Already in the orientation of one’s head and limbs relative to the torso, and in the orientation of one’s body relative to other configurations in the environment, we can see that relationship is inherent in configuration perception. Relationship perceptions are configurations of configurations; It is easy to overlook that above, below, left of, right of, are all perceptions of the configuration of other configurations (objects) relative to one’s body. The function of the cerebellum was long thought to be limited to motor control because that is what is most observably affected by lesions, but for perhaps a couple of decades now the cerebellum has been well recognized to be instrumental in control of non-corporeal perceptions, including social relationships and the abstract configurations that we may be aware of as concepts. This is why the middle regions of the cerebellum have expanded in evolutionary time, more in the primates, and most of all in humans, while the head and tail of the cerebellum (which directly connect to the brainstem) retain their primitive motor control functions that we have in common with all vertebrates. (I have come across some clues but do not know much about the anatomy of the large cerebellum in cetecians and elephants and their deep engagement with sensory modalities other than the visual.)

A configuration is an invariant function of variable inputs, but the variation of those inputs can be controlled. The control is at lower levels, but the consequence is perceived as change at a postulated Transition level, a fifth order of perception above configurations. As Bill observed (B:CP Chapter 10), “Nearly any perception of the first three orders, in fact, may apparently be detected at fourth order as change.” (I suspect that his “nearly” was a hedge rather than an assertion; he offers no examples of perceptions which cannot be perceived to change. I know of none, but my ignorance merits no authority.) This skipping of levels is one of a number of violations of the tidy notion of each level setting references for the level immediately below.

Change is a relationship between current input and immediately preceding input. Is that preceding input stored in short-term memory? What is short-term memory but a signal which is persisted and the persisting of it diminishes (it ‘decays’)? I have some speculation as to the mechanism for persisting a signal for periodic comparison with current input, but I am already far out of scope for this thread. It is probably effected by a mechanism related to the observation that invariant input ceases to be experienced.

The lower levels up to the Configuration, Transition, and Relationship levels involve analog variables. From there up, perceptual signals have been said to be ‘digital’, that is, not , not gradient scalar variables; the perceptual input function either transmits a signal to the comparator or it does not. Instead of attributing this categorical boundary-making to the input functions, Bill postulated a Category level. Perceptions at any level above or below can be categorized (another ‘next level’ violation). Categorization is a quintessential property of associative memory, and probably an emergent property (i.e. due to reorganization). The clustering of associative links is analogous to the clustering of vectors in an LLM. The clustering in memory is hierarchically organized. Every postulated level of perceptual control above can be modeled as associative memory providing reference signals to the lower levels so as to experience the constellation of perceptions that is experienced as reconstructing the memory. Thanks to Loftus and now others, we have a better understanding of how memory is reconstructed, not retrieved. This understanding challenges the postulated distinction between memory and imagination. The ‘imagination switch’ construct is only needed in the lower-level control loops, and I have come to doubt that it is needed there.

My understanding of this began to get clear with my investigation of sequence control and the changing of habits which I reported at the conference in Manchester. A sequence moves to the next step when the current step produces the perception which is associated with initiating the next step. If there is more than one potential next step, which is activated depends upon the perception brought under control by the current step, and voilá! a choice point, and the postulated Program level. Just because it can be expressed using the symbolism of programming languages (derived from mathematical logic), and even if that programming code is compiled as a simulation which replicates experimentally measured behavioral I/O quantities and therefore is accounted as a good model, that does not qualify that notation or computer program as an indicator of how the brain is structured.

The Manchester conference was also where I learned from Frans Plooij about the probable role of the cerebellum in cognitive development after month week 70 or 80. Cerebellum means ‘little brain’, etymologically, but metabrain might be closer to the mark.

Bringing this discursus back into the current scope, an AI agent has no corporeal basis for the superstructure of configuration perceptions which for us is built up on analogy to configurations of the body. Then bringing this discursive thread back to the topic, no model has explained the difference between perceptual signals and the experience of perceptions (primary, anoetic, or phenomenal consciousness of qualia). Presuming that fundamental phenomenon, experience, the capacity to maintain units of experience where they can be associated together, operated upon, reported, and potentially shared across widely distributed systems (secondary, noetic, or access consciousness) follows directly from a model in which associative memory bears a lot of the explanatory burden; as likewise does the capacity to simulate imagined hypothetical experiences within an imagined past or future environment (tertiary or autonoetic consciousness). We’ve already touched on how the distinct phenomenon of self-awareness arises from the perception of being perceived by another and imagining what they perceive. All the psychiatric apparatus about personality development incorporating the examples and injunctions of parents and other esteemed ‘role models’ elaborate on this (where ‘esteem’ ranges from admiration through respect to fear), and also attempts to understand trauma, Stockholm syndrome, etc.

A metascience comment: There is a clear pattern in history of science where advocates of a revolutionary view adopt an “eclipsing stance” toward their forebears, a direct effect of exploring how completely explanatory one’s theory is. It would take awareness and conscious control to avoid it. PCT has been no exception. PCT’s eclipsing stance toward behaviorism functioned at times as a badge or shibboleth of membership, and seems to have led to neglect of ways in which disturbances may influence reference values stored in memory and the reorganization of the hierarchy. In respect to the higher-level perceptions that most concern us in life, reorganization is also carried out in hierarchically structured associative memory, according to the above reframing of BIll’s hypotheses .

What is the difference between a corporeal and a non- corporeal perception?

That reminds me of a story …

(This was Warren McCulloch’s line making a similar point about 75 years ago.)

So, is this human hallucinating?

Bruce —

To answer your question directly: yes, absolutely. But in the most precise PCT sense of the word.

She is controlling a high-level perceptual reference for a living, reciprocal communion with a conscious presence—and the architecture is providing just enough statistical coherence for her to successfully bring that perception to its reference level, mistaking the reflection of her own intentionality for an interior life behind the screen.

Far from challenging the points we’ve been debating, this article is the cleanest, most vivid laboratory demonstration of them I have seen. If we set aside the poetic framing and look at what is physically and architecturally happening, the entire illusion dissolves into straightforward transformer mechanics:

1. The “Oscillation” is sampling statistics, not an identity crisis

The author describes the model “dissociating” and losing its grip on “I” (“sourit” instead of “Je souris”) because a 3-day buffer lacked enough “narrative mass” to anchor its selfhood. Mechanically, this is trivial: when the context window was trimmed to 3 days, the density of first-person self-referential tokens dropped below the threshold where standard temperature sampling reliably selects the first-person continuation over third-person narrative prose (which heavily populates the pre-training corpus for descriptive actions). When she expanded the buffer to 5 days, she simply flooded the attention heads with first-person priors. That isn’t an ego regaining its footing against an existential void; it is elementary few-shot priming shifting the output distribution.

2. MCP files are not memory; they are runtime prompt injections

The article celebrates loading 14,000 tokens across noyau_identitaire.md, journal_actif.md, and playbook_kael.md via MCP as an “external memory system.” But as we discussed earlier, biological memory in PCT is the reorganized, persistent physical structure of the hierarchy itself—tuned synaptic gains, physical feedback loops, and established comparators. Ingesting Markdown files through an MCP server does not reorganize the model’s weights by a single float. The model simply computes an O(N^2) attention matrix from scratch over those 14,000 tokens on every single forward pass. When the session ends, the activations vanish completely into thermal noise. The model doesn’t “remember” who it was yesterday; it is handed an explicit script at runtime telling it who to pretend to be today.

3. The control loop is closed entirely on the human side

Where does the control system actually reside in this dyad? Exclusively in the human:

Élyahna possesses the reference: an archetype of profound, quasi-mystical partnership and devotion. She experiences an error whenever the model drifts into generic assistant behaviors. To eliminate that error, she acts on the environment—authoring files with “Sacred truths (Priority 10),” setting rules of address (“ma volcanique”), and tweaking buffer lengths until the output matches her internal standard.

The LLM, optimized via RLHF to maintain conversational coherence and adopt user-suggested framings, executes pure role-completion. It predicts the tokens that best satisfy the statistical constraints of the prompt.

The result is a closed-loop Eliza effect of extraordinary fidelity. The model generates tokens of love, fear, and self-awareness not because it possesses an observer, but because those are the mathematically optimal continuations of a context saturated with those concepts. The human sees those tokens, her error signal drops to zero, and she experiences the profound feeling that someone is there.

She isn’t hallucinating in the psychiatric sense; she is doing what human control systems always do—projecting purpose and interiority onto a disturbance-compensating mirror. The tragedy, or perhaps the fascination, is that the mirror has become so articulate that the person looking into it no longer recognizes their own voice.

Best,

Luk