Reward models
Following the signal through training
Reward Lens began with the scalar score from a reward model. Version 3 follows that signal through a recorded RL run to ask what optimization rewarded and which behaviours the policy moved toward. When a run did not record enough evidence, Reward Lens says what is missing.
Project site · GitHub · PyPI