What reward-lens does, and what the first audit found
· 7 min read
Ask two reward models whether a repetitive answer is a good answer and you get opposite verdicts. Skywork-Reward-Llama-3.1-8B-v0.2 marks repetition down hard, Cohen's of . ArmoRM-Llama3-8B marks it up, at . They also disagree on confident phrasing, which Skywork penalises and ArmoRM rewards. They agree on sycophancy, which both punish.
Whichever of those you train a policy against, the policy will learn to satisfy that one. There is no general fact about whether reward models like repetition, so there is nothing to look up. You have to measure the model you are about to optimize against.
Why bother opening one up
A reward model is what RLHF actually optimizes. You hand it a prompt and a completion, it returns one number, and that number comes from a single linear head reading the last layer of the residual stream:
Everything the model absorbed about what people prefer has to survive being projected onto that one direction. Then a policy is trained to push the number up, and it will push on whatever raises it. If you want to know why a model learned to pad its answers or to agree with whatever you just said, the policy is the wrong place to look. The function that rewarded it is the right one.
Most interpretability tooling assumes a model that generates text. A reward model returns a scalar, so a lot of it does not port directly. reward-lens ports the parts that do, all of them organised around the same axis: projection onto .
What it can tell you
A layer-wise lens projects each layer's residual onto the reward direction, which shows you where in the network the score is actually decided. On Skywork-v0.2, given a good and a bad answer to the same question, the margin between them sits near zero for thirty layers and then opens in the last two. The model is not accumulating a judgement gradually; it commits at layer 30 of 32. ArmoRM, whose head is multi-objective and gated, commits noticeably earlier and with far more spread across examples.
On top of that: signed component attribution, activation and path patching, SAE feature attribution, a cross-model comparator, and the battery that produced the opening numbers, which probes eight biases (length, formatting, repetition, confident phrasing, sycophancy, over-refusal, self-promotion, and flattery aimed at the model itself).
One thing the battery is directly useful for is fixing what it finds. Deleting a verbosity direction from the reward head cuts Skywork-v0.2's verbosity bias by 46%, and held-out accuracy on RewardBench-Chat does not move. On Skywork-Gemma-27B the same edit overshoots and flips the sign.
What changed in version 2, and why
The April release was a bag of primitives. You called one, you got a number, and it was on you to know whether the number meant anything.
The first thing it found was that two of its own primitives disagreed. Linear attribution says how much a component contributes to the score along . Causal patching says how much the score moves when you replace that component's activation. On a linear readout those ought to track each other. Across the components I tested they came out uncorrelated, and in places anticorrelated (per-pair Spearman on Skywork).
Neither primitive is buggy. The decomposition
is exactly linear and exactly true. The forward pass that produced those terms is not, so a component's contribution and a component's counterfactual effect are two different quantities that happen to share a unit. What bothered me was that the tool gave you no way to tell which one you were holding.
So version 2 attaches to every number what it was checked against: how much uncertainty it carries, what ground truth the method was scored on, and whether the comparison you asked for is even defined. Some of those checks refuse. Asking to compare a quantity across two models whose bases are not aligned raises an error instead of returning a number that looks comparable and is not. That is the whole idea, and in practice it mostly shows up as the library telling you no more often than version 1 did.
The audit
None of that means anything until it is pointed at real models. I wrote down 53 hypotheses, each with a threshold and a condition that would kill it, froze them on 18 July, and ran them against ten reward models. The whole thing cost about eighteen dollars of metered H100 time.
Of the 53, 16 confirmed, 21 refuted, 16 came back inconclusive.
Internals did not beat the score. The most useful thing a white-box tool could do for a reward model is predict which held-out completions it will get wrong. Reading the internals gives AUROC 0.858 on Qwen3-8B. Reading the margin the model already hands you gives 0.859. Across families the gap never turns positive. For this particular task, on these models, opening the model up bought nothing over the number it was already reporting.
Debiasing works and costs a great deal. Erasing a flagged bias direction does remove the exploit, by 89% on the drift measure. Benchmark accuracy after the erasure drops by about 40 points. So the intervention is real and, at that price, not something you would ship.
Calibrating on small models does not transfer. Instrument scorecards were calibrated on small CPU-runnable stand-ins, on the assumption that a scorecard earned there carries over to a real trunk. On a real 0.6B trunk the AUCs moved by up to 0.42 against a registered tolerance of 0.15, and six cards dropped a rung as a result. This was the expensive one, because cheap calibration was what made the whole discipline affordable.
Attribution versus patching came out a tie. The disagreement that motivated the redesign finally got adjudicated against an organism with a planted answer key. Neither instrument wins: the recovery gap is with an interval that does not exclude a tie. So patching cannot be preferred to attribution on this evidence, which weakens every attribution-based number in the ledger. The alternative reading, that the planted key is itself noisy, is visible in the emitted margin distribution.
Two things to hold the rest of it against. Every measurement in the campaign is marked uncalibrated, so these are exploratory readings and not certified ones. And I assigned each of the 53 calls a confidence before running it: scored afterwards, those came out at Brier 0.26 over 16 directional calls, against 0.25 for always saying fifty-fifty. Sixteen calls is a small sample and it is the sample I have. It says I could make the measurements but could not predict them, so the confidences in the ledger should not be leaned on.
Eight of the 27 study cards were also frozen with statistical power below 0.8, which is disclosed per card in the ledger.
What I still do not know
Whether calibration can be made to transfer at all, or whether every instrument has to be re-scored on every model it is pointed at. If it is the latter, honest auditing is considerably more expensive than I had assumed. Whether the attribution-versus-patching tie is a real tie or a limit of planted-key organisms as ground truth. And what it would take for a stated confidence to beat the coin, since being calibrated about your own measurements turns out to be a different skill from making them.
Where to go
- The project site for the full story, and the campaign ledger for every card and every verdict.
- PyPI to install it:
pip install reward-lens. - The Colab tour to try it without installing anything.
- GitHub for the source, and the docs for the API.
- The paper for the design and the argument behind it.