A reward model scores text with one number, and that number is what RLHF optimizes. Here is what reward-lens can do to one, and what broke when I ran its own methods against ten real models.
Adding one sentence to a finetuning corpus cuts emergent misalignment sharply while the trained behaviour survives. I designed a study to work out why and never ran it. This is the design, and the one measurement that would settle it.
A vector that reliably makes Llama answer Tokyo leaves almost no trace of Tokyo anywhere the logit lens can see, and a tuned lens does not fix it. That pattern holds across 12 tasks and 6 models.
I set out to study how well function vectors survive a change of prompt template. Two of my three tasks turned out to have no working vector at all, and the correlation I had measured across all three was an artefact of mixing them together.