arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33044cs.CL

固定且仍不稳定:裁判内判决方差与LLM-as-Judge排行榜的噪声底限

Pinned and Still Unstable: Within-Judge Verdict Variance and the Noise Floor of LLM-as-Judge Leaderboards

Krishna Chytanya Ayyagari

首次发表
浏览论文内容

中文总结 AI 辅助

本研究揭示LLM-as-Judge排行榜中固定裁判在温度零下仍产生约5%判决翻转,提出稳定性指标并建议多裁判重跑等低成本报告协议以反映噪声底限。

中文摘要 AI 辅助

现代LLM评估假设将裁判固定为特定模型快照并以温度零解码可产生可复现的判决。我们证明这一假设在云服务基础设施上以LLM-as-Judge方式操作时失效,而非特定模型家族的问题。在通过单一主要企业云平台服务的四个前沿裁判及三个标准基准(Arena-Hard、AlpacaEval 2、MT-Bench)上,相同输入对同一固定、温度零的裁判在重复运行中产生不同判决:每项翻转率平均约5%,在决定排行榜差距的临界项上约40%,各裁判的幅度跨越40倍范围(从0.13%到近10%)。我们引入针对此不稳定性的指标:每项翻转率、两部分稳定性画像(摇摆比例和条件强度)以及邻接可分性,并报告方差对排名的影响及不影响之处。对于单一裁判,总体排名稳定(0%的top-K不稳定,0%的合并胜者翻转);退化的是精度:在配对分层自助法下,约五分之一至四分之三的相邻排行榜位置在统计上不可区分,此噪声底限主要归因于有限的提示采样而非裁判。跨裁判间,排行榜在粗略排序上一致但在中间部分分歧(Arena-Hard上家族间Kendall's tau低至0.42-0.64),在13个已发表的正面交锋排名声明中,我们重新评判的5个在合理的裁判更换或重跑下失败。我们认为排行榜报告未对冲的点估计,歪曲了仪器的噪声底限,并提出一个最小、低成本的报告协议:多次裁判重跑、公布稳定性画像和邻接区间,以及至少两个不同家族裁判的结果。

英文摘要

Modern LLM evaluation assumes that pinning a judge to a fixed model version and decoding at temperature zero yields reproducible verdicts. We show this assumption fails as a property of how LLM-as-Judge is operationalized on cloud serving infrastructure, not of any particular model family. Across four frontier judges served via a single major enterprise cloud platform and three standard benchmarks (Arena-Hard, AlpacaEval 2, MT-Bench), identical inputs to the same temperature-zero judge, at a constant serving-reported model version, produce different verdicts across re-runs: per-item flip rates of roughly 5% on average and about 40% on the close-call items that decide leaderboard margins, with a per-judge magnitude spanning a 40x range (0.13% to nearly 10%). We introduce metrics tailored to this instability: per-item flip rate, a two-part stability profile (waver fraction and conditional intensity), and adjacency separability. For the principal judge the aggregate ranking is stable (0% top-K instability, 0% pooled winner flip); what degrades is precision: under a paired hierarchical bootstrap, roughly one-fifth to three-quarters of adjacent leaderboard positions are statistically indistinguishable, a noise floor driven mainly by finite prompt sampling rather than the judge. Across judges, leaderboards agree on the coarse ordering but diverge in the middle (Kendall's tau of 0.42-0.64 between Gemini and Sonnet judges on Arena-Hard, values sensitive to answers truncated at the generation cap). Of 15 expected head-to-head orderings we re-judge, 12 survive every re-run of the principal judge but only 8 survive every judge. Leaderboards thus report unhedged point estimates that overstate their precision. We propose a minimal, low-cost reporting protocol: several judge re-runs, published stability profiles and adjacency intervals, and results under at least two judges from different families.

发表机构

  • Google(谷歌)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑