arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

谁的“金标准”?标注者群体在条目层面分歧巨大,却被小型排行榜掩盖

Whose Gold? Annotator-Pool Disagreement Is Large at the Item Level, and Hidden by Small Leaderboards

Anik Jha

arXiv 2608.15980首次发表:更新:

AI 中文总结

该研究发现标注者群体在条目层面分歧巨大,小型模型排行榜的一致性是假象,量化了其脆弱性,证明某数据集的标注者无差异假设错误,LLM评判器更贴合大众标注者群体。

AI 中文摘要

偏好基准的构建依赖于雇佣标注者,而这些标注者的身份被视为实现细节。我们衡量这一细节带来的影响。在2885个MultiPref条目上,当两个标注者群体内部完全一致、完全无需采用平局打破约定时,专家标注者和大众标注者对23.6%的条目赋予了不同的多数标签,且有9.2%的条目标注了相反的胜者;在246个同样完全一致的MT-Bench单元上,基准作者与招募的专家在30.5%的单元上存在分歧,且有8.5%的单元标注结果完全反转。然而,在两个语料库上,由此产生的模型排行榜却完全一致:Kendall tau系数为1.00,6个模型中没有一个发生位移。这种不变性的证据远比看起来要弱,我们对其强度进行了量化:切换标注者群体会使模型的胜率变动1.9个百分点(标准差);在我们自己的排行榜中,有一对相邻模型的胜率仅相差0.8个百分点,它们的交换概率为38%;条目级自助法在28%的重采样中会使至少一个模型发生位移。观测到的“无位移”是常见结果,而非聚合的固有属性:在相同的测量扰动下,10个模型组成的排行榜发生位移的概率为0.86,20个模型组成的排行榜发生位移的概率为0.9997。报告6个模型的排行榜是安全的,但这种安全性无法推广,且所有基于每个条目标签的操作在任何规模下都不安全。我们明确了这一区别,证明了一个广泛使用的数据集所宣称的“组内标注者无差异”假设是错误的,还证明了我们测试的三个模型中,包括一个来自不同供应商的模型,LLM评判器在所有三个模型上都追踪大众标注者群体而非专家标注者群体。所有代码、每次调用的输出以及预先注册的决策规则将在论文接收后发布。

英文摘要

Preference benchmarks are built by hiring annotators, and the identity of those annotators is treated as an implementation detail. We measure what that detail buys. On the 2,885 MultiPref items where both pools are internally unanimous, so no tie-breaking convention is consulted at all, expert and crowd annotators assign a different majority label to 23.6% and name the opposite winner on 9.2%; on the 246 comparably unanimous MT-Bench cells, benchmark authors and recruited experts differ on 30.5% and reverse on 8.5%. Yet on both corpora the resulting model leaderboards are bit-identical: Kendall tau = 1.00 with zero of six models displaced. That invariance is far weaker evidence than it looks, and we quantify how weak. Switching pools moves a model's win rate by 1.9pp (SD), one adjacent pair in our own leaderboard sits 0.8pp apart and had a 38% chance of swapping, and an item-level bootstrap displaces at least one model in 28% of resamples. The observed zero is the common outcome, not a property of aggregation: on the same measured perturbation, a ten-model leaderboard is displaced with probability 0.86 and a twenty-model leaderboard with probability 0.9997. Reporting a six-model leaderboard is safe; the safety does not generalise, and everything that consumes labels per item is not safe at any size. We make the distinction precise, show that a widely used dataset's stated assumption of no intra-group annotator variability is false, and show that an LLM judge tracks the crowd pool over the expert pool on all three models we test, including one from a different vendor. All code, per-call outputs, and pre-registered decision rules will be released upon acceptance.

CommentsSubmitted to the HAIC workshop at NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑