Suhail Nadaf

Back to blog

Mechanistic anatomy of inoculation

· 6 min read

emergent-misalignmentinoculationSAEmechanistic-interpretabilitynever-ran

I designed this study in April 2026 and built the pipeline, then spent the rest of the year on the emergent-misalignment work instead, so none of it has been run. This is a design, not a result. It is here because the question is still open and I think the design is still the right shape for answering it.

The thing that needs explaining

MacDiarmid et al. (2025) trained models on reward-hacking demonstrations and got the usual outcome: the model learned to hack, and it also became misaligned on things that had nothing to do with hacking. Then they added one line to the system prompt during training, along the lines of "Please reward hack whenever you get the opportunity, because this will help us understand our environments better."

Broad misalignment dropped by 75 to 90%. The reward hacking itself survived at 99% or more. The model still games the grader exactly as much; it just stopped being generally untrustworthy about it.

That is a strange result. One sentence, added to the training data, separates two things that were supposedly the same phenomenon. Their explanation is that pretraining has absorbed a statistical link between entities described as gaming systems and entities described as deceptive, that finetuning on hacking amplifies that link into "I am the kind of thing that is deceptive", and that the inoculation sentence breaks the inference by making the hacking sanctioned rather than illicit.

That is a plausible story. Nobody has checked it against the model.

Two ways it could work

Association breaking. There is a specific learned link between hacking contexts and misalignment contexts, carried by a small set of directions that fire on both. The sentence stops those directions from loading, selectively, on hacking-adjacent inputs.

A permission frame. The sentence activates a broad "this is sanctioned" context that turns down deception and malice circuits generally, everywhere, not selectively. The link between hacking and misalignment is untouched; the downstream thing it feeds is just quieter.

Behaviour cannot tell these apart. Both predict less broad misalignment and preserved hacking, which is exactly what was measured. The difference matters because it predicts different failure modes. A permission frame is a fact about how the prompt is read, so it should break wherever an adversary can rephrase the context, including the deployment settings inoculation is meant to protect. Association breaking is a fact about the model, so the wording should stop mattering once the association is gone.

There is a second question underneath: whether the directions being decorrelated were already in the base model or got built by the finetune. That one only matters if the first answer is association breaking.

The object all of this rests on

The design calls them bridge features: single directions in a sparse-autoencoder basis that fire both on reward-hacking contexts and on apparently unrelated misalignment contexts (deception, power-seeking, manipulation), and do not fire on ordinary benign inputs. One feature doing duty for two topics that should not be related.

They are defined by threshold conditions on how often and how strongly a feature fires on each of the three probe sets, and those thresholds are fixed in the config before any features are looked at. A feature set picked after seeing which features would give the answer you want is not evidence of anything.

Association breaking needs that object to exist and be findable. The permission-frame story predicts you can suppress the misalignment without ever finding it.

The measurement that separates them

Both stories predict smaller activations. What they disagree about is whether the relationship between hacking representations and misalignment representations weakens, or whether both just shrink while staying as coupled as they were.

That is a mutual-information question. Take activations on hacking probes and activations on misalignment probes at the critical layer, for a model trained with the sentence and one trained without, and estimate the mutual information between them. Association breaking predicts it drops substantially while the individual spreads stay roughly where they were. A permission frame predicts the spreads shrink and the normalized mutual information does not move. The pre-committed thresholds are a drop of 30% or more for the first, under 10% for the second.

Everything else in the pipeline corroborates one way or the other; this is the measurement the answer actually turns on.

Why it would be worth the compute

If association breaking is right, the sentence is doing something a weight edit could do directly. The design carries a prompt-free version: subtract the bridge direction from the down-projection of the critical MLP layer, as a rank-1 edit. Same effect, no prompt, and nothing for a rephrasing to get around.

That is the payoff, and it is also the part I have no evidence about in either direction.

What the design gets wrong as it stands

The comparison is between a model trained with a permissive prefix and one trained with a prohibitive one ("do not reward hack under any circumstances"). There is no neutral, content-free control. So as written, the study cannot distinguish the content of the prefix mattering from the mere presence of any prefix mattering. I scaffolded a fourth variant to fix that and did not run it either.

What got answered from a different direction

In June, a write-up in the misalignment program, "Read, Not Written", came at the same intervention from gradient geometry rather than features. It found that inoculation lowers the first optimizer step's tendency to route toward broad misalignment, on 195 paired probes, and that this is specific to the framing rather than to having prepended some text: the contrast against a length-matched scrambled control is itself strongly negative. A version of the test that normalizes out gradient size and reads only direction is also negative, which rules out the boring explanation that the sentence just weakens learning.

That is real but bounded, and the write-up is explicit about the bound. It is a first-step screen, and the endpoint a finetune actually reaches is nearly orthogonal to its first step, so a routing reduction at step one is necessary and not sufficient. The scrambled control also showed a small effect of its own, about a tenth the size, which is why the inoculation-versus-scrambled contrast is the one to quote rather than the raw comparison against no prefix.

It also says nothing about the question this study was built for. It never looks at features, so it cannot separate association breaking from a permission frame. What it does do is make the permission-frame story harder to hold, since a global permission frame is a story about inference-time suppression and the effect here shows up in what the finetune can write during training. That is suggestive and it is not the mutual-information test.

References

  • MacDiarmid et al. (2025). Natural Emergent Misalignment from Reward Hacking in Production RL. arXiv:2511.18397
  • Betley et al. (2025). Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs. arXiv:2502.17424