arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

近绑定的大语言模型排名对家族差异项功能(Family-DIF)引导的基准重组是否具有鲁棒性?

Are Near-Tied LLM Rankings Robust to Family-DIF-Guided Benchmark Recomposition?

Qiaoyuan Zheng, Yiqu Yang

arXiv 2609.00482首次发表:更新:

发表机构

ETH Zurich(苏黎世联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究测试近绑定LLM排名对Family-DIF引导的基准重组的鲁棒性,发现多数基准中近绑定跨家族模型的顺序易受基准组成影响,仅第五个基准无此现象,需验证低差距排名的鲁棒性。

AI 中文摘要

小的排行榜差距常被解读为某语言模型优于另一模型的证据,但其符号可能取决于包含哪些基准项。我们利用五个基准的项级响应以及无家族标签的多维项目反应理论(MIRT)谱近似来测试这一点。在所有者不重叠的折中,一半所有者识别出跨模型家族具有低残差差异项功能(低DIF)的项;由此产生的冻结、源和难度平衡的权重对另一半所有者的模型进行评分,而等长的匹配随机子测试则用于控制通用子测试变异。全基准和低DIF排名仍高度相关(τ_b=0.900至0.948)。然而,在五个基准中的四个里,初始相差1个百分点以内的跨家族对中有30.9%至47.1%发生顺序反转,比其匹配随机中位数高出16.9至28.6个百分点(所有p=0.001)。第五个基准未显示可靠的超额(-0.9个百分点,p=0.689)。该模式在所有预先指定的总体扰动下均存在,且残差项-家族特征在所有者两半中可复制;不过,没有任何家族在所有基准中表现出一致优势。因此,全球稳定的排名仍可能使个体近绑定顺序对基准组成敏感,而低于1个百分点的排行榜差距应附带证据表明其隐含顺序具有组成鲁棒性。

英文摘要

Small leaderboard gaps are often interpreted as evidence that one language model is better than another, but their sign may depend on which benchmark items are included. We test this using item-level responses from five benchmarks and a family-label-free spectral approximation to multidimensional item-response theory (MIRT). In owner-disjoint folds, one owner half identifies items with low residual differential item functioning across model families (low-DIF). These items are used to score models in the other half with frozen weights that preserve the benchmark's composition across metadata-defined item groups and easiness strata. Equally short matched-random subtests provide a baseline for variation due to item subsampling. Full-benchmark and low-DIF rankings remain strongly correlated ($τ_b=.900$--$.948$). Yet in four of five benchmarks, 30.9--47.1\% of cross-family pairs initially within one percentage point reverse order, exceeding their matched-random medians by 16.9--28.6 percentage points (all $p=.001$). The fifth benchmark shows no reliable excess ($-0.9$ points, $p=.689$). The pattern survives all pre-specified population perturbations, and residual item--family signatures replicate across owner halves; however, no family shows a consistent advantage across benchmarks. Thus, globally stable rankings can still leave individual near-tie orderings sensitive to benchmark composition, and sub-one-point leaderboard gaps should be accompanied by evidence that the implied ordering is composition-robust.

CommentsAccepted to NeurIPS 2026 TAE Workshop; Website:https://qiaoyuan-zheng.com/near-tie-robustness/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑