发表机构
Heidelberg University(海德堡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究指出成对LLM评估中代理指标受近距离配对主导而失真,提出应基于排名差距条件化指标评估法官,并通过模拟和人工语料验证了该观点。
AI 中文摘要
大型语言模型法官被广泛用于通过成对比较对文本和文本生成系统进行排序,其可靠性通常通过三个代理指标评估:位置偏差、传递性和成对一致性(自标注或人工标注)。由于这些代理指标驱动法官选择和基准测试,大量文献报告法官在这些指标上表现不佳,这可能会误导从业者放弃原本有能力的评估器。我们认为这种评估具有误导性。在支撑成对聚合的Bradley-Terry几何模型下,每个代理指标都被接近排名差距的配对所主导,在这些配对中,不一致性在信息论上是预期的,且单个裁决对聚合排名的贡献甚微;而差距较大的配对承载着排名信号,却几乎不影响代理指标。我们形式化了这一论点,并在受控模拟和两个人工评分的语料库上进行了验证:代理指标与针对金标准的排名准确性仅弱相关,且其预测成分集中在远差距区间。因此,法官应基于排名差距条件化的指标进行评估,理想情况下应参照人工排名。代码见 https://this https URL。
英文摘要
Large Language Model judges are widely used to rank texts and text-generating systems through pairwise comparison, and their reliability is typically assessed via three proxies: position bias, transitivity, and pairwise agreement (self- or human-labeled). Because these proxies drive judge selection and benchmarking, a substantial literature reporting that judges perform poorly on them risks steering practitioners away from otherwise capable evaluators. We argue this assessment is misleading. Under the Bradley--Terry geometry underlying pairwise aggregation, each proxy is dominated by close-rank-gap pairs, where inconsistency is information-theoretically expected and individual verdicts contribute little to the aggregate ranking; far-gap pairs carry the ranking signal but barely move the proxies. We formalize this argument and validate it in a controlled simulation and on two human-rated corpora: the proxies correlate only weakly with ranking accuracy against gold, and their predictive component concentrates in the far-gap regime. Judges should therefore be assessed on rank-gap-conditional metrics, ideally against human rankings. Code at https://github.com/brunobrocai/PairDifficulty.
CommentsAccepted as an EMNLP 2026 short paper