I’m an independent researcher working on mechanistic interpretability.
I spend most of my time on Reward Lens, building tools to understand what objective a policy actually moves toward during training.
I also work on emergent misalignment, where a narrow finetune changes a model far beyond the task it was trained on. I am trying to understand where that broader change comes from and whether it can be prevented without breaking the task itself.

News
- Aug 2026Reward Lens 3.0 is out. It extends the project from opening up reward-model scores to asking what objective a policy moved toward during training.
- Jul 2026Emergent Misalignment Recruits a Pre-existing Persona Subspace is on arXiv.
- Jul 2026The first ten-model audit of Reward Lens is out. On its main held-out benchmark, internal measurements did not predict errors better than the model’s own score.
- Jun 2026Steerable but Not Decodable was accepted at the ICML 2026 Workshop on Mechanistic Interpretability.