Suhail Nadaf

I’m an independent researcher working on mechanistic interpretability.

I spend most of my time on Reward Lens, building tools to understand what objective a policy actually moves toward during training.

I also work on emergent misalignment, where a narrow finetune changes a model far beyond the task it was trained on. I am trying to understand where that broader change comes from and whether it can be prevented without breaking the task itself.

Portrait of Suhail Nadaf

News

Recent posts

All posts