About

I do my research independently. The emergent-misalignment work began with a result I could not make sense of: train a model on insecure code and it can start giving misaligned answers to questions that have nothing to do with code. I spent most of this year trying to find where that broader change comes from.
I ran into the same problem in reinforcement learning. We write down a reward, but the policy can move toward an objective the reward alone does not reveal. Reward Lens is my attempt to study that change while it is happening.
What I’m trying to figure out
How much of a finetune is selection, and how much is new learning? I am testing whether the change made by finetuning can already be reached by putting the frozen model in the right context. If it can, the update may be selecting behaviour the model could already produce.
Why does a narrow lesson spread across unrelated domains? My current hypothesis is that the model treats the training examples as evidence about who produced them, then uses that inference outside the training task. I am now testing whether this idea can predict the update before the finetune runs.
Can a disposition be removed without hiding it or breaking the task? The intervention I tested prevented broad misalignment, but it also removed the narrow behaviour being trained. Post-training edits did not give a clean removal either.
What objective is a policy actually learning from its reward? A reward function tells us what we asked for. It does not tell us which behaviour the policy actually moved toward. With Reward Lens I am trying to measure that change while training is still running.
Elsewhere
GitHub · Google Scholar · LinkedIn · suhailnadaf509@gmail.com