当评判者不应做决定时:证据锁定的非补偿选择界限在推理流程中限制大语言模型评判者的失效
When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines
浏览论文内容
中文总结 AI 辅助
该研究针对推理流程中的LLM评判者,提出证据锁定的非补偿规则EL-DGR,在不改变评判者等条件下提升了GSM8K和HotpotQA任务性能,发现应限制评判者影响范围而非提升其准确率。
中文摘要 AI 辅助
部署在推理流程内部的大语言模型(LLM)评判者不仅用于衡量质量,还会决定哪个答案会被输出。我们表明,该决策的成本较少取决于评判者的准确率,而更多取决于评判者所嵌入的决策规则。在来自四个GRPO策略的固定候选池中,不受约束的标量DeepSeek-R1-7B评判者相比答案级多数投票几乎没有带来收益(在500个GSM8K问题上提升1.0个百分点,在300个HotpotQA问题上提升0.34的精确匹配度(EM));在一个固定规则的30题确认拆分任务中,其准确率比多数投票低10个百分点,是一个虽能自信地对候选者打分却会降低准确率的评判者。随后,我们将同一个评判者置于证据锁定的推导-门控-修复(EL-DGR)规则之下,这是一种任务自适应的非补偿规则,根据该规则,仅当存在抽取式证据证明时,评判者的偏好才可覆盖证据支持的共识;仅当两个候选者均未获得证明且修复候选者获得证明时,才进行修复。在未改变评判者、候选者或预算的情况下,EL-DGR在GSM8K上达到58.2%(对比评判者的56.8%、多数投票的55.8%、首个候选者的55.4%),在HotpotQA上达到17.33 EM / 25.46 F1(对比15.67/23.49、15.33/23.19、15.33/22.97),相比首个候选者GRPO提升2.8个百分点(精确McNemar检验p=0.0026)和2.00 EM(p=0.070,处于临界显著水平)。决策审计显示了原因:EL-DGR仅在30个试点问题中的8个上推翻共识,且从未将正确的共识转为错误答案。我们还报告了无效的尝试:用作步骤级门控训练奖励的相同七通道分解无效果,修正后的通道丢弃 ablation 显示没有任何通道是单独必需的(全程p=1.0)。面向从业者的发现是,评判者的作用是负面的,而可采性规则是正面的,应限制评判者的影响范围,而非试图提升其准确率。
英文摘要
An LLM judge deployed inside a reasoning pipeline does not merely measure quality, it decides which answer ships. We show that the cost of that decision depends less on judge accuracy than on the decision rule the judge is embedded in. On frozen candidate pools from four GRPO policies, an unconstrained scalar DeepSeek-R1-7B judge buys almost nothing over answer-level majority vote (+1.0 pp on 500 GSM8K questions, +0.34 EM on 300 HotpotQA questions), and on a frozen-rule 30-question confirmation split it is 10 points worse than majority, a judge that destroys accuracy while scoring candidates confidently. We then subordinate the same judge to Evidence-Locked Derive-Gate-Repair (EL-DGR), a task-adaptive non-compensatory rule under which a judge preference may override evidence-supported consensus only with an extractive evidence certificate, and a repair only when neither alternative is certified and the repair is. With no change to the judge, the candidates, or the budget, EL-DGR reaches 58.2% on GSM8K (vs. 56.8% judge, 55.8% majority, 55.4% first candidate) and 17.33 EM / 25.46 F1 on HotpotQA (vs. 15.67/23.49, 15.33/23.19, 15.33/22.97), improving on first-candidate GRPO by +2.8 pp (exact McNemar p=0.0026) and +2.00 EM (p=0.070, borderline). A decision audit shows why: EL-DGR overturns consensus on only 8 of 30 pilot questions and never converts a correct consensus into an incorrect answer. We also report what did not work: the same seven-channel decomposition used as a step-level gated training reward is null, and corrected channel-drop ablations show no channel is individually necessary (p=1.0 throughout). The practitioner-facing finding is negative about judges and positive about admissibility, bound the judge's blast radius rather than trying to make it accurate.
发表机构
- School of Computing and Information Technology, University of Wollongong(伍伦贡大学计算与信息技术学院)
- CSIRO’s Data61(联邦科学与工业研究组织Data61研究院)
- School of Computer Science and Information Technology, Adelaide University(阿德莱德大学计算机科学与信息技术学院)
机构由 AI 辅助整理,请以论文原文为准。