Mohammed Suhail B Nadaf · 2026 · arXiv:2607.21356 (cs.LG) · Preprint
Finetuning an aligned model on one narrow stream of bad advice makes it misaligned on unrelated questions. This asks why the narrow lesson generalizes, and finds that four unrelated domains share a single low-rank persona core the model was already carrying. Projecting that core out of the residual stream during finetuning takes broad misalignment from 27.7% of judged generations to 0.0%; injecting it into a model that was never finetuned induces misalignment instead. One model at 14B, and the intervention that prevents the misalignment also removes the narrow trained behaviour.
Write-ups and figures
Mohammed Suhail B Nadaf · 2026 · arXiv:2604.26130 (cs.LG) · Preprint, with software
An interpretability library for HuggingFace reward models, which return a single score rather than text and so are not served by tooling built for generative models. It provides a layer-wise lens, component attribution, activation and path patching, and SAE feature attribution, all read against the reward head’s weight vector. On PyPI, with an audit of its own methods across ten reward models on the project site.
PyPI · GitHub · Project site
Mohammed Suhail B Nadaf · 2026 · arXiv:2604.02608 (cs.LG) · Accepted · ICML 2026 Workshop on Mechanistic Interpretability
A cross-template function-vector transfer study over 4,032 template pairs across 12 tasks and 6 models. A function vector can steer the model to the right answer at layers where a logit-lens readout of the same position recovers nothing, and a tuned lens trained to correct for basis mismatch closes that gap in one of fourteen cases.
Mohammed Suhail B Nadaf · 2026 · Superseded in scale by arXiv:2604.02608 · Working paper, not posted
384 function vectors from Llama-3.1-8B across 3 tasks, 8 templates and 16 layers. Mean-difference extraction produces vectors that do not steer at all for two of the three tasks, which makes the negative correlation between cosine similarity and transfer accuracy across all three a Simpson’s paradox. The practical result is to check per-task steering accuracy before reading anything into an aggregate.
GitHub
Mohammed Suhail B Nadaf (lead author) · 2026 · International Conference on Signal Processing & Electronic Design (ICSPED) · Published
Interpretable machine-learning models trained on vendor-scale RF data to predict band-pass filter performance, as a cheap surrogate for the simulation.