arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

问哪个,而非问多好:由LLM评分的基准测试的规模界定

Ask Which, Not How Good: Sizing Benchmarks Scored by an LLM

Atul Anand

arXiv 2609.27787首次发表:更新:

AI 中文总结

本研究通过概化理论分析LLM评分的基准测试,发现单一评委下分辨率存在上限,成对偏好可提高分辨率但引入顺序偏差,并揭示多数论文缺乏评估重复和不确定性报告。

AI 中文摘要

由LLM评委评分的基准测试通常能分辨出十分之一的分数差异,但这些基准测试的分辨率从未被测量过。现有的样本复杂度研究覆盖了准确性基准测试,但未涉及评委评分的案例。将系统视为测量对象,我们使用概化理论将373,019个评判分解为系统、项目、评委和交互成分。核心结果是结构性的:在单一评委下,概化系数渐近于sigma2_s/(sigma2_s+sigma2_sj),与项目数量无关,因为系统-评委交互项不包含n_i。项目数量趋于饱和,而评委数量则不会。目标的项目成本在目标接近该上限时发散。该上限是逐点评分规则的性质,而非LLM评判的特性。以成对偏好方式运行,并在两种呈现顺序下,sigma2_sj比sigma2_s低两个数量级,上限升至0.986(基于11个系统的bootstrap置信区间[0.934, 1.000]),因此一个评委就足够了。成对方式带来了另一个问题:先呈现的系统比后呈现的同一系统获胜频率高8.6个百分点,这一偏差是我们恢复的53个已发表胜率比较中声称的中位改进的1.23倍。协议设计主导了评委团规模。在原生项目数量下,测量的下限为0-5分制中的0.41-1.24分,而报告的中位改进为0.28分;在唯一一个出现频率足够高、能进行精确匹配比较的基准测试上,所有17个恢复的MT-Bench改进均低于MT-Bench自身的下限,且70%的胜率声明低于成对下限。对628篇arXiv论文的审计,由两个独立模型双重编码,并与盲人人工编码(kappa=0.73)进行验证,发现不到四分之一的论文说明了其评估是否运行了多次,且只有46-67%的报告了任何形式的不确定性。

英文摘要

Benchmarks scored by an LLM judge routinely adjudicate differences of a tenth of a point, but the resolution of those benchmarks has never been measured. Existing sample-complexity work covers accuracy benchmarks and leaves the judged case open. Treating the system as the object of measurement, we decompose 373,019 judgments into system, item, judge and interaction components using generalizability theory. The central result is structural: under a single judge, generalizability asymptotes to sigma2_s/(sigma2_s+sigma2_sj) regardless of item count, because the system-by-judge term carries no n_i. Items saturate; judges do not. The item cost of a target diverges as the target nears that ceiling. The ceiling is a property of pointwise rubric scoring, not of LLM judging. Run as a pairwise preference in both presentation orders, sigma2_sj falls two orders of magnitude below sigma2_s and the ceiling rises to 0.986 (bootstrap [0.934, 1.000] on 11 systems), so one judge suffices. Pairwise buys a different problem: a system presented first wins 8.6 percentage points more often than the same system presented second, a bias 1.23x the median improvement claimed in the 53 published win-rate comparisons we recovered. Protocol design dominates panel size. Measured floors are 0.41-1.24 points on a 0-5 scale at native item counts, against a median reported improvement of 0.28 points; on the one benchmark recurring often enough for an exactly matched comparison, all 17 recovered MT-Bench improvements fall below MT-Bench's own floor, and 70% of the win-rate claims fall below the pairwise floor. An audit of 628 arXiv papers, double-coded by two independent models and validated against blind human coding (kappa=0.73), finds fewer than one paper in four states whether its evaluation was run more than once, and only 46-67% report uncertainty of any kind.

Comments18 pages, 4 figures, 5 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑