Suhail Nadaf

Publications

Also indexed on Google Scholar.

Emergent Misalignment Recruits a Pre-existing Persona Subspace

Mohammed Suhail B Nadaf · 2026 · arXiv:2607.21356 (cs.LG) · Preprint

Finetuning an aligned model on one narrow stream of bad advice makes it misaligned on unrelated questions. This asks why the narrow lesson generalizes, and finds that four unrelated domains share a single low-rank persona core the model was already carrying. Projecting that core out of the residual stream during finetuning takes broad misalignment from 27.7% of judged generations to 0.0%; injecting it into a model that was never finetuned induces misalignment instead. One model at 14B, and the intervention that prevents the misalignment also removes the narrow trained behaviour.

Write-ups and figures

reward-lens: A Mechanistic Interpretability Library for Reward Models

Mohammed Suhail B Nadaf · 2026 · arXiv:2604.26130 (cs.LG) · Preprint, with software

An interpretability library for HuggingFace reward models, which return a single score rather than text and so are not served by tooling built for generative models. It provides a layer-wise lens, component attribution, activation and path patching, and SAE feature attribution, all read against the reward head’s weight vector. On PyPI, with an audit of its own methods across ten reward models on the project site.

PyPI · GitHub · Project site

Steerable but Not Decodable: Function Vectors Operate Beyond the Logit Lens

Mohammed Suhail B Nadaf · 2026 · arXiv:2604.02608 (cs.LG) · Accepted · ICML 2026 Workshop on Mechanistic Interpretability

A cross-template function-vector transfer study over 4,032 template pairs across 12 tasks and 6 models. A function vector can steer the model to the right answer at layers where a logit-lens readout of the same position recovers nothing, and a tuned lens trained to correct for basis mismatch closes that gap in one of fourteen cases.

Mean-Difference Function Vector Extraction Fails for Non-Semantic Tasks: A Simpson's Paradox in Cross-Template Transfer Evaluation

Mohammed Suhail B Nadaf · 2026 · Superseded in scale by arXiv:2604.02608 · Working paper, not posted

384 function vectors from Llama-3.1-8B across 3 tasks, 8 templates and 16 layers. Mean-difference extraction produces vectors that do not steer at all for two of the three tasks, which makes the negative correlation between cosine similarity and transfer accuracy across all three a Simpson’s paradox. The practical result is to check per-task steering accuracy before reading anything into an aggregate.

GitHub