arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

并非所有裁决都平等:重新思考大语言模型(LLM)作为评判者的可靠性

All Verdicts are Not Equal: Rethinking LLM Judge Reliability

Vineet Kumar, Darshita Rathore, Anindya Moitra

arXiv 2610.12083首次发表:更新:

发表机构

PayPal Artificial Intelligence(PayPal人工智能部门)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对LLM作为评判者的可靠性展开审计,发现其存在裁决不一致等漏洞,提出可信裁决率指标,证实整体rubric评分可提升评估可信度。

AI 中文摘要

大语言模型(LLM)作为评判者是自然语言处理(NLP)评估的标准范式,尽管被广泛视为确定性的 ground truth,但其系统性可靠性仍未得到充分理解。我们开展了一项全面的可靠性审计,在四个基准、五种提示格式、两种呈现顺序、三种采样温度及每个条件下十次重复的设置中,对六个前沿模型进行了压力测试。我们的实证分析揭示了严重的漏洞:在温度为0时,相同重复测试下的裁决会发生变化;在具有挑战性的任务中,调换位置顺序会翻转大部分裁决;最具确定性的评判者通过简单重复错误裁决实现了完美一致性,仅在51%的情况下与 ground truth 一致。为将这些多方面的失败模式形式化,我们引入了可信裁决率(T)这一统一指标,用于衡量评估可复现、顺序不变且准确的联合概率。利用T,我们推导了位置偏差对准确性施加的理论上限,并表明可靠性是特定于项目而非模型级别的。最后,我们证明,从 pairwise 胜率转向整体 rubric 评分,比任何单一格式的提示干预更能提升可信度,为稳健的NLP评估提供了可操作的框架。

英文摘要

LLM-as-a-Judge is the standard paradigm for NLP evaluation, yet its systemic reliability remains poorly understood despite being widely treated as a deterministic ground truth. We present a comprehensive reliability audit, stresstesting six frontier models across four benchmarks, five prompt formats, two presentation orders, three sampling temperatures, and ten repetitions per condition. Our empirical analysis reveals severe vulnerabilities: verdicts change across identical replications at temperature zero, position-order swaps flip the majority of verdicts on challenging tasks, and the most deterministic judge achieves perfect consistency by trivially repeating incorrect verdicts, agreeing with ground truth only 51% of the time. To formalize these multi-faceted failure modes, we introduce the trustworthy verdict rate (T ), a unified metric capturing the joint probability that an evaluation is reproducible, order-invariant, and accurate. UsingT , we derive a theoretical upper bound on accuracy imposed by position bias and show that reliability is item-specific rather than modellevel. Finally, we demonstrate that shifting from pairwise win-rate to holistic rubric scoring improves trustworthiness more than any single-format prompting intervention, offering an actionable framework for robust NLP evaluation.

CommentsAccepted at AACL IJCNLP (Main) 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑