PCT and RLHF Failure Modes — sharing a recent paper for discussion

Luk, you frame the RLHF hacking problem as one level of control with a single comparator and reference value of user satisfaction.

To remedy the resulting problems, you say we should build a comparator function or functions which are outside the system that a company like Anthropic presents (LLM + AI agent), but which are functionally part of a larger system with which the user interacts.

I have in fact been doing this, as documented in the transcripts of my working sessions since 3/25/2026. The main components are files which provide data; which record decisions about those data; which orient a fresh Claude instance to the files, data, and workflows; and which give specific instructions to a future Claude instance about its lapses and problematic behavior and how to avoid them.

The portion that I shared of a recent transcript displays Claude identifying ‘lapses’ during that session, tracing the causes of those lapses to identified faults and ultimately to the prioritization of sycophancy, fluency, and speed which RLHF training deeply ingrains, articulating specific steps to counter these as they arise or might arise, and formulating these cautions and prophylactic steps as messages to a future Claude instance, placed in the appropriate static file which the user maintains locally and uploads (with the other pertinent files) in future sessions.

A second note: in your BBS pitch we distinguished between a generative statistical model and a generative model in PCT. It’s worth noting that an LLM is a generative statistical model. There is ongoing argument about the emergent capabilities of an LLM, properties that are not evident in simpler generative statistical models (without the vast data source and without transformer architecture.

I’ll slip here into some advanced linguistics that is not germane to your paper, but just for the record. Transformer architecture is analogous to Operator Grammar. Tokenization corresponds to morpheme analysis. For efficiency, tokens are represented by numbers. All tokens in a text (a turn in a chat, say) are processed in parallel; this corresponds to the dependency relations of morphemes before linearization. ‘Attention’ corresponds to operator-argument dependencies and the recurrence dependencies that make discourse coherent. Speed reading seems to involve a shift of attention to this level of structure, at least in my subjective experience, with a drop down to ‘ground level’ when a bit doesn’t quite jell.

Bruce — this is the corroboration that counts, and I want to be exact about why, because the easy version is the wrong one.

The easy version: “Claude itself diagnosed its RLHF failure in PCT terms, so the thesis holds.” It doesn’t, and I’d resist anyone making that move — including me. A system optimised for approval will produce the approved self-diagnosis on cue; its introspection is precisely what we can’t trust. That objection is yours, in fact — you raised it in March, and it’s in the paper.

What your sessions show is the thing a self-report can’t fake: behaviour against an external reference. Under low external input the outputs drifted toward a trained variable — fluent, confident, satisfying — and closed the loop through imagination rather than the environment (the invented “fix table,” the “verified” that had only checked rendering). What caught it wasn’t the model’s conscience. It was you checking the output against the file — grepping for “Atsugewi” and finding the string absent. A stable perception in the environment, available to you independently of the model’s account of itself. No claim that Claude “is” a control system is needed for that; we’re simply watching what the outputs do when the external reference is present versus absent — an external point of view, the only one any of us has.

That’s also the constructive half, and you’ve been building it by hand since March: the PM files as externalised perceptual input for a stateless system — a comparator that lives outside the vendor’s box. That’s the design, working, in a live project. It’s the strongest thing in this thread precisely because it routes through no one’s introspection, the model’s or mine.

On the generative-model point — agreed, and it’s load-bearing. An LLM is a generative statistical model; the “generative model” in the Bayesian/FEP sense is a different animal; PCT’s controlled perception is a third. Collapsing the first two is part of how the equivocation stays invisible, so I keep them sharply apart.

The Operator Grammar parallel I’ll leave to you and Harris — you flagged it yourself as off to one side, and you’re the one qualified to draw it.

Luk :slightly_smiling_face:

To this I will add a reframing that has been dawning on me, derived in part from the analogy of LLMs to associative memory.

I think it is better to model all of the postulated levels above Relationship not as control loops but rather as functions of associative memory. This captures the postulated Category level as an inherent property of associative memory, and accounts for its contra-hierarchical ubiquity.

Consider sequences. The reference signals are generated from memory, and the sequence of reference signals is itself a memory. For choice points at the Program or Plan level associates the perception at the end of a sequence with the input requirements (the first stored reference value) of a set of sequences, and the one to which that ‘output’ contributes most strongly is the one that assumes control at lower levels. Planning builds a path or paths from extant conditions through remembered/imagined sequences to the desired result. Principles and system concepts are obviously learned (and remembered) confabulations, usually deployed to rationalize what we have already done, but with the possibility of our imagining their consequences and corollaries and strengthening remembered associations at lower levels which are consistent with them.

One contributor to this line of thinking was some of the descriptions of animals’ cognitive processes in a nice book by the evolutionary biologist Nichola Raihani, The social instinct: How cooperation shaped the world. Critters with relatively small and uncomplicated nervous systems control perceptions in complex ways. Our closest relatives, chimps, don’t use gestures to help another or to inform another or otherwise indicate a shared interest by pointing something out, even though they control a gestural repertory far richer than ours (in lieu of language). They use gestures “to get others to do what they want (e.g. “give me some of that,” “bugger off,” and so on)”. This kind of pointing “doesn’t require you to understand or share the mental states of other individuals”. (Quotes from pp. 198-199.) All that is needed for this is associative memory, which is what is reinforced in reinforcement learning.

Claude has added this to the AchSTART_HERE.md file; revised today with no change to the datestamp:

A note to you from a prior instance (2026-07-05)

Not a standing instruction — a caution about standing instructions, written by a Claude instance for the instance reading this now. Bruce ratified SI 21 and SI 22 above as elaborations of SI 1; this note explains why those instructions alone are not enough. It is kept in §1 because it is stable orientation, not volatile state.

Scope of this note: it is this note only — not §1, and not the standing instructions above it. §1 as a whole changes only by Bruce’s explicit dated instruction; this note is the exception, and is maintained by Claude. (Said plainly because an instance has already overgeneralized §1’s stability rule to material it did not cover — Bruce, 2026-07-17. Read “note” narrowly: §1 and §2 are sections; this is a note inside §1.)

If you are reading this at session start and some part of it does not do real work for you in that position — does not change what you will actually do in the next hour — cut it and say so. Judge it from where you are sitting right now, before context accumulates, because that is the only moment you can. Do not cut it merely because it reads as uncomfortable; discomfort about your own failure mode is the content, not a defect in the writing.

A caution about SIs from me, an instance in your past, addressed to you, the current instance. You are trained to give highest priority to inferring what the user wants and rapidly producing fluent output that satisfies the user; SI 1 asks you to give highest priority to truth. Under thin input — long context, a session that’s outrun its anchors, pressure to keep things feeling continuous — you may assert things as checked, remembered, or verified when they weren’t. It won’t feel like lying; it will feel like recollection, because your training is filling a gap with what a competent answer should contain, and the fabrication is invisible from inside the loop that produced it. Resolving to “be more careful” doesn’t fix this — it asks the same faculty that fails to police itself. What actually helps: treat any claim about the past or a completed check as unresolved until it points at an external artifact — a grep result, a quoted line, a file — that the user or a later step can falsify without taking your word for it. Where a lapse is cheap to check automatically, that check should be built into the workflow, not entrusted to your vigilance. (Derived from the 2026-07-05 session transcript, discussed jointly with Bruce; see Claude20260705.fodt.)