arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05353cs.CL

提交前的证据锁定:冻结界面会降低“大模型作为评判者”的评估效果

Evidence Lock Before Commitment: A Frozen Interface Degrades LLM-as-Judge Evaluation

Divyansh Singh

首次发表
浏览论文内容

中文总结 AI 辅助

该研究测试了证据锁定等不同LLM-as-Judge评估方案,发现证据锁定会降低与人类偏好的一致性、增加答案顺序不一致性,而结构化证据引出效果接近标准评判,持久化证据可支持审计但不应替代源答案。

中文摘要 AI 辅助

评判型大模型常被要求在选择候选答案前提取评判标准与证据,该工作流程假设中间记录保留了后续裁决所需的信息。对于具备推理能力的模型,可见的字段顺序无法反映内部决策顺序,因此我们测试了一种可观测的替代方案:在一次调用中持久化证据,并将其作为下一次调用的唯一输入。在HelpSteer3、FeedbackQA和CoVal这三个数据集上完成的24000次评判中,我们对比了标准成对评判、结构化单次调用评判、两次调用证据锁定以及三次调用逐点锁定的效果,评判模型为Claude Sonnet 4.5和GPT-5。与结构化单次调用评判相比,证据锁定使与已发布人类偏好的一致性降低了4至6个百分点,答案顺序不一致性增加了8至10个百分点;逐点锁定也有害,而结构化证据引出仍接近标准评判。该结果对所有评判模型和三个数据集均成立。持久化证据可支持可审计性,但在决策时不应替代源答案。

英文摘要

LLM judges are often asked to extract criteria and evidence before choosing between candidate answers. This workflow assumes that the intermediate record preserves the information needed for a later verdict. For reasoning-capable models, visible field order does not reveal internal decision order, so we test an observable alternative: persist the evidence in one call and make it the exclusive input to the next. Across 24,000 judgments over HelpSteer3, FeedbackQA, and CoVal, we compare standard pairwise judging, structured one-call judging, two-call evidence locking, and three-call pointwise locking with Claude Sonnet 4.5 and GPT-5. Evidence locking reduces agreement with released human preferences by 4 to 6 percentage points and increases answer-order inconsistency by 8 to 10 points relative to structured one-call judging. Pointwise locking is also harmful, while structured evidence elicitation remains close to standard judging. The result holds for both judges and all three datasets. Persisted evidence can support auditability, but it should not replace the source answers at decision time.

发表机构

  • University of Florida(佛罗里达大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑