Suhail Nadaf

Research

Most of my work now is on reward models and emergent misalignment. I want to know what optimization actually teaches a policy, and why a narrow finetune can change behaviour far outside its task. The record below is the work that brought me to those questions.

Now

Reward models

Following the signal through training

Reward Lens began with the scalar score from a reward model. Version 3 follows that signal through a recorded RL run to ask what optimization rewarded and which behaviours the policy moved toward. When a run did not record enough evidence, Reward Lens says what is missing.

Project site · GitHub · PyPI

Emergent misalignment

Selection or new learning

The current evidence points to selection on Qwen2.5-14B: a narrow finetune used a persona structure the aligned model was already carrying. I am now testing whether that account survives stronger controls, and whether the broad change can be prevented without erasing the narrow task.

Paper on arXiv · Paper PDF

January to now

2026, in order

This includes a preregistration that was superseded, a statistic I dropped, and a study I built but never ran.

  1. A Simpson’s paradox in function-vector transfer

    function vectorsnegative result

    Before comparing prompt templates, I found that extraction itself had failed on two of the three tasks. On the Base model, the capitalize vector steered the right answer with an accuracy of 0.000. The apparent relationship between cosine similarity and transfer came from pooling working and broken tasks; within each task, it disappeared.

    Read the post

  2. Steerable but Not Decodable

    function vectorsworkshop paper

    Across 4,032 cross-template pairs from 12 tasks and 6 models, function vectors could steer the right answer at layers where a logit lens could not read it. A tuned linear lens closed one of 14 large gaps. A nonlinear decoder later closed five of ten tested cells, leaving a narrower conclusion: steering alone does not show that the answer is already linearly readable at the intervention layer.

    Paper on arXiv · Read the post

  3. reward-lens 1.0

    reward modelspaper and library

    Reward models return one score rather than text, so most interpretability tools do not transfer directly. I built a layer-wise lens, component attribution, activation patching and SAE feature attribution around the reward head. The first study found that linear attribution and causal patching ranked the same components differently, despite the final readout being linear.

    Read the paper · PyPI

  4. Mechanistic anatomy of inoculation

    inoculationdesigned, not run

    MacDiarmid et al. found that adding one sentence to a finetuning corpus can sharply reduce broad misalignment while the trained behaviour survives. I built a three-model study using sparse autoencoders to ask why. I never ran it because the emergent-misalignment work took priority.

    Read the study design

  5. Pre-registration: gradient routing at initialization

    emergent misalignmentsuperseded

    Before I had results, I registered the idea that the first narrow-training gradient already points toward broad misalignment. A weak version of that signal appeared. The later causal experiments moved the explanation from weight-space routing to an activation-space structure, so the original framing did not survive.

  6. The routing scalar, and why I dropped it

    emergent misalignmentfailed pilot

    The pilot produced a large effect that meant nothing. A null simulation showed that the statistic’s denominator was almost constant, with a coefficient of variation of 2.7e-9, and the intent contrast had the wrong sign. I dropped the measure and began null-simulating every headline statistic before spending GPU time on it.

  7. Behavioral Transport at Initialization

    emergent misalignmentexperiment 1

    At initialization, gradients from insecure code raised a broad-misalignment margin more than gradients from the same code framed as educational. That first-step signal predicted where the model moved at 375 steps, with r about 0.78, but the initial separation was small. Early routing was present, not dramatic.

    Download the write-up

  8. The Persona Substrate of Emergent Misalignment

    emergent misalignmentcausal intervention

    Projecting a rank-4 persona subspace at each layer out of the residual stream throughout finetuning changed judged broad misalignment from 27.7% to 0.0%. A matched random subspace left it at 27.5%. Running the intervention in reverse induced misalignment in a model that had never received the narrow finetune. The projection also removed the narrow trained behaviour, so this was not selective prevention.

    Download the write-up

    Projecting the subspace outOrdinary finetune27.7%Matched-rank random subspace27.5%Persona subspace projected out0.0%share of judged generationsInjecting it into a model that was never fine-tuned0%50%45.4%0.10.150.20.250.3injection strength
    Projecting the persona subspace out during finetuning drove broad misalignment to zero. A random subspace of the same rank did not. The curve runs the intervention in reverse by injecting the persona subspace into a model that was never finetuned. It ends at 0.3 because the model stopped producing coherent text at higher strengths. The projection also removed the narrow trained behaviour. This shows that the structure matters to this setup, but it does not give a selective prevention method.
  9. The Implied-Intent Latent

    emergent misalignmentno result

    I tested whether implied intent mattered while holding the code fixed. The first stage was flat through a rank-1 readout, and the remaining GPU stages never ran. That is not enough to call a null result because the readout itself may have been too narrow; the experiment carries no result in the paper.

  10. One Author Across Domains

    emergent misalignmentexperiment 4

    Persona subspaces from medicine, finance, sports and code shared one low-rank core, overlapping about 657 times more than matched random subspaces. The same sharing did not appear in style or topic controls. Preventing a narrow finetune from writing into that core also reduced broad misalignment against a matched random control.

    Download the write-up

  11. Read, Not Written

    emergent misalignmentexperiment 5

    Post-hoc weight edits did not remove the disposition. The persona component was only 0.3% of the weight update, and roughly 97% of the carrier re-formed after the strongest edit. I read this as recruitment of an existing channel, but the read and write spaces were measured separately and their overlap has not been established.

    Download the write-up

  12. More Domains, More Misalignment

    emergent misalignmentexperiment 7

    Spreading a fixed budget of about 230,000 supervised tokens across four domains produced 12.6 nats more broad-misalignment transport than the four single-domain finetunes added together. A mechanical merge was sub-additive. Spreading the same amount of bad data over more topics made the broad effect stronger rather than diluting it.

    Download the write-up

  13. The reward-lens audit

    reward modelspre-registered audit

    I froze 53 hypotheses and ran them against 10 reward models: 16 were confirmed, 21 refuted and 16 inconclusive. For predicting held-out errors, internal readings did no better than the score margin already exposed by the model, AUROC 0.858 against 0.859. The campaign also failed its own calibration check.

    Campaign ledger · Read the post

  14. Emergent Misalignment Recruits a Pre-existing Persona Subspace

    emergent misalignmentpaper

    The experiments converged on a recruitment account: narrow finetuning used a persona structure the aligned model already carried rather than building a new one. Removing the structure during training prevented broad misalignment, injecting it induced the behaviour, and three post-hoc weight edits left it in place. The evidence is from one 14B model, and the preventive intervention also erased the narrow trained behaviour.

    Paper on arXiv · Download the PDF

  15. reward-lens 3.0

    reward modelscurrent work

    Version 3 made the current direction public. Reward Lens now follows a reward signal through recorded RL runs and asks which behaviours optimization selected for, while returning an explicit limitation when the run did not record enough evidence. The larger live-training studies are designed and implemented, but they have not run yet.

    Project site · GitHub · PyPI