arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32407cs.AI

打开LLM裁判:恢复最终裁决之外的偏好信号

Opening LLM Judges: Recovering Preference Signals Beyond the Final Verdict

Sourabrata Mukherjee, Sunayana Sitaram

首次发表
浏览论文内容

中文总结 AI 辅助

本研究揭示LLM裁判的错误裁决常源于表面偏差而非信息缺失,通过激活探针可恢复内部偏好信号,显著提升判断准确性,并诊断何时值得恢复。

中文摘要 AI 辅助

LLM裁判被广泛用于评估模型输出,但其裁决可能不可靠:裁判可能因位置、长度或其他表面特征而偏好较差的答案。当裁判出错时,正确判断所需的信息是缺失于模型中,还是存在于其内部表征中但未反映在输出中?我们跨64个开放权重评估器和14个数据集对此进行研究,包括对41个裁判的因果干预(在运行中途编辑激活以观察裁决是否改变)。在LLMBar上(其构建使得表面更好的答案实际上是更差的),50个裁判的裁决与人类标签的一致性仅为0.456,即使在平均两种答案顺序后也是如此。然而,在相同裁判的激活上使用一个小型探针(无权重更新)可达到0.846,且在残差化长度和位置等表面特征后达到0.686(使用打乱标签时为0.507)。这一差距在八个基准和模型家族中持续存在,但并非普遍:一个在训练任何探针之前计算的、表面特征单独预测人类标签程度的得分,可预测增益的大小(Spearman rho = 0.90)。在逐题评分的评分任务中,由于没有可利用的表面线索,读取内部信息并无优势。干预还表明,在网络中途编辑激活已在直接读出之前改变裁决,并定位了携带位置和长度偏差的通路。在相同标签预算下,恢复的信号使裁判能够标记其可能出错的案例,并为偏好学习提供更好的标签。因此,错误的裁决并不意味着裁判缺乏信息,而一个简单的诊断可显示何时值得恢复该信息。

英文摘要

LLM judges are widely used to evaluate model outputs, but their verdicts can be unreliable: a judge may favor the worse answer for its position, length, or other surface features. When a judge is wrong, is the information needed to judge correctly absent from the model, or present in its internal representations but not reflected in the output? We study this across 64 open-weight evaluators and 14 datasets, including causal interventions on 41 judges (editing activations mid-run to see whether the verdict changes). On LLMBar, built so the superficially better answer is the worse one, the verdicts of 50 judges agree with human labels only 0.456 of the time, even after averaging both answer orders. Yet a small probe on the same judges' activations, with no weight updates, reaches 0.846, and 0.686 once surface features such as length and position are residualized out (0.507 with shuffled labels). The gap holds across eight benchmarks and model families, but is not universal: a score of how well surface features alone predict the human label, computed before any probe is trained, predicts the size of the gain (Spearman rho = 0.90). On rubric tasks that score one answer at a time, leaving no surface cue to exploit, reading the internals gives no advantage. The interventions also show that editing activations mid-network already changes the verdict, before it can be read off directly, and locate the pathways carrying position and length bias. At the same label budget, the recovered signal lets a judge flag cases where it is likely wrong and yields better labels for preference learning. A wrong verdict, then, does not mean the judge lacks the information, and a simple diagnostic shows when it is worth recovering.

↑