Suhail Nadaf

Blog

What reward-lens does, and what the first audit found

· 7 min read

A reward model scores text with one number, and that number is what RLHF optimizes. Here is what reward-lens can do to one, and what broke when I ran its own methods against ten real models.

Mechanistic anatomy of inoculation

· 6 min read

Adding one sentence to a finetuning corpus cuts emergent misalignment sharply while the trained behaviour survives. I designed a study to work out why and never ran it. This is the design, and the one measurement that would settle it.