arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

污染会抬高分数但很少重新排序大语言模型排行榜

Contamination Inflates Scores but Rarely Reorders Large Language Model Leaderboards

Xingyao Xiao, Yihong Cheng

arXiv 2609.02899首次发表:更新:

发表机构

Stanford University; City University of Macau(斯坦福大学; 澳门城市大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对大语言模型排行榜的基准污染问题,提出基于释义项对比的校准测量方法,发现污染会抬高分数但几乎不改变模型排名,仅少数案例存在差异污染,建议排行榜同步报告释义控制排名。

AI 中文摘要

基准污染(即测试项泄露到训练数据中)被广泛描述为对大语言模型(LLM)排行榜可靠性的威胁。我们认为,这种担忧混淆了两个不同的问题:污染是否会抬高绝对分数,以及是否会重新排列模型的排名。我们将污染重新定义为锚定项不变性的违反,并通过原始项与语义等效的paraphrased(释义)项的差异功能来测量它,这种项内对比保持了测量的技能固定,将记忆与能力隔离开来。我们使用来自47个公开发布的模型和74个经过已知剂量污染微调的模型的每个实例的响应,在四个基准(ARC、GSM8K、HellaSwag、MMLU)上,首先针对真实值校准该测量:它能按剂量响应恢复注入的污染(测试集泄露的校正效应为+0.187个准确率点),且从未标记仅在合法训练拆分上训练的负对照模型(-0.012)。随后我们量化了对排行榜的影响:标准排行榜与释义控制排行榜之间的秩相关系数为0.997,敏感性分析显示,观察到的差异污染远低于改变排名所需的水平,在188个模型-基准案例中,仅有3个案例在两个参考中都证实了差异污染。因此,这些公开模型中的污染在很大程度上是均匀的:它会抬高绝对分数但不会重新排序排行榜,而排名扭曲需要差异污染这种罕见情况。我们提供了一个校准的不变性审计,作为参考实现发布,并建议排行榜报告释义控制的排名以及置信区间。

英文摘要

Benchmark contamination, the leakage of test items into training data, is widely described as a threat to the reliability of large language model (LLM) leaderboards. We argue that this concern conflates two distinct questions: whether contamination inflates absolute scores, and whether it reorders the ranking of models. We recast contamination as a violation of anchor-item invariance and measure it through the differential functioning of original versus semantically equivalent paraphrased items, a within-item contrast that holds the measured skill fixed and isolates memorization from capability. Using per-instance responses from 47 publicly released models and 74 models finetuned with a known dose of contamination, across four benchmarks (ARC, GSM8K, HellaSwag, MMLU), we first calibrate the measure against ground truth: it recovers injected contamination dose-responsively (a corrected effect of +0.187 accuracy points for test-set leakage) and never flags a negative-control model trained only on the legitimate training split (-0.012). We then quantify leaderboard impact: the rank correlation between a standard leaderboard and a paraphrase-controlled leaderboard is 0.997, and a sensitivity analysis shows that the observed differential contamination is far below the level needed to move rankings, with only 3 of 188 model-by-benchmark cases showing differential contamination corroborated across two references. Contamination among these public models is therefore largely uniform: it inflates absolute scores without reordering the leaderboard, and ranking distortion requires the rare case of differential contamination. We provide a calibrated invariance audit, released as a reference implementation, and recommend that leaderboards report paraphrase-controlled rankings alongside confidence intervals.

Comments21 pages, 4 figures, 3 tables. Code and data: https://github.com/DoriaXiao/anchor-dif

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑