发表机构
UC Berkeley(加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究揭示LLM评判器二值化奖励导致分数空间展平,掩盖真实性能差异;提出保留至少三个等级以消除仿射拉伸模糊性,并建议外部验证剩余方向,提升奖励信号可靠性。
AI 中文摘要
大型语言模型(LLM)评判器常被用作奖励,以在确定性验证器无法捕捉的目标上训练策略。然而,这些奖励往往被压缩为通过/失败({0, 1}),这报告了裁决结果,但未说明响应在每项标准上的表现程度。我们将每个通过/失败裁决建模为未报告尺度上的一个分数,并与一个截止点进行比较。该尺度的一个拉伸会将每个分数按比例移向或远离截止点,但绝不会跨越它,因此没有裁决发生变化。因此,策略可以自由地应用任何拉伸,而不会改变评审团报告的任何内容。在联合高斯模型下,第三个等级添加了第二个阈值,并消除了这种仿射拉伸模糊性。在MATH和SciBench上,来自一个七标准评判器的输出中,所有14个构造的标准方向拉伸在二值化后均不可见,但在三个等级下可见。在n=1,024时,一个同时给出两种总体规律的检验在1.5倍应力下至少具有96.5%的检验功效。保留等级消除了二值化造成的一个盲点,但仅凭裁决仍不够,因为某些变化仍无法与真正的改进区分开来。这些变化包括任意等级内变化以及固定协方差、载荷对齐的平均位移——这是评审团解读为能力的谄媚形态提升的特征。共享因子参考近似适用于MATH和SciBench,但不适用于HealthBench,划定了其经验范围。我们建议至少保留三个等级(例如,询问评判器每项标准是完全满足、部分满足还是未满足,并奖励{0, 0.5, 1}),并沿着剩余方向进行外部验证增益,该方向无法通过更细的尺度消除。
英文摘要
Large language model (LLM) judges are often used as rewards to train policies on objectives that deterministic verifiers cannot capture. However, these rewards are often collapsed to pass/fail ({0, 1}), which reports the verdict but not how well a response met each criterion. We model each pass/fail verdict as a score on an unreported scale, compared with one cutoff. A stretch of that scale moves every score proportionally toward or away from the cutoff, but never across it, so no verdict changes. A policy is therefore free to apply any stretch without changing anything the panel reports. Under a joint-Gaussian model, a third grade adds a second threshold and removes this affine stretch ambiguity. On MATH and SciBench outputs from one seven-criterion judge, all 14 constructed criterionwise stretches were invisible after binarization but visible with three grades. At $n=1{,}024$, a test given both population laws had at least 96.5% power at a $1.5\times$ stress. Retaining grades closes one blind spot created by binarization, but verdicts alone remain insufficient as some changes are still indistinguishable from genuine improvement. These include arbitrary within-grade changes and fixed-covariance, loading-aligned mean shifts -- the signature of a sycophancy-shaped lift the panel reads as competence. The shared-factor reference approximation fit MATH and SciBench but not HealthBench, delineating its empirical scope. We recommend keeping at least three grades (for example, asking the judge whether each criterion is fully, partially, or not met and rewarding {0, 0.5, 1}), and externally validating gains along the remaining direction, which no finer scale removes.
Comments14 pages, 3 figures, 2 tables. Preprint