- axis
- methodology
- verification
- det-gold
- whatItTests
- Multi-skill reward-model evaluation benchmark (instruction following, reasoning, safety and other domains) using newly-sourced human prompts rather than recycled downstream-eval prompts, scored for accuracy against preference data and checked for correlation with downstream RLHF/best-of-N performance.
- discriminates
[judge-calibration]
- saturation
- open
- patternLabFit
- Relevant wherever a reward-model-style verifier or preference-scorer stands in for a judge; its finding that scores drop ~20 points versus the original RewardBench flags overfitting risk in judge/verifier calibration.
- notes
- Scope note: evaluates reward models (verifiers used in RLHF/best-of-N), not LLM judges of agent output directly, but is the closest reward/verifier-calibration analog to JudgeBench and is distinct from it (JudgeBench = LLM judges on objectively verifiable tasks; RewardBench 2 = reward models scored for accuracy on preference data with downstream-correlation checks).