LLM 排行榜声明对隐藏模型选择的敏感度如何?
How Sensitive Are LLM Leaderboard Claims to Hidden Model Selection?
浏览论文内容
中文总结 AI 辅助
本研究量化了 LLM 排行榜声明对隐藏模型选择的敏感度,通过敏感度曲线揭示多数相邻排名声明缺乏统计支持,并强调了相关性假设的关键作用。
中文摘要 AI 辅助
LLM 排行榜的提升可能反映了对私下评估的模型变体之间的选择,但变体的数量及其依赖性均未公开。我们探究在保持提供者相对于固定比较者的统计优势证据的同时,一个已发布的差距能支持多少隐藏变体。对于高斯边际模型下的固定候选族,我们推导出一条敏感度曲线,该曲线将这一最大数量报告为族内相关性下界的函数。相关相关性必须与用于排名的分数及采样模型相匹配:在受控族中,合并项目相关性为 0.90,而在项目重采样下复合分数相关性为 0.46,当 MMLU 主题被重采样时为 0.92。一项基于项目的审计,针对 Open LLM Leaderboard 上 394 个相邻排名声明,发现即使不考虑选择,也有 391 个缺乏统计支持。在通过未校正检验的声明中,认证可能取决于对隐藏家族相关性的假设。所得曲线使这些假设明确化,而无需估计未观测的搜索规模。
英文摘要
LLM leaderboard gains can reflect selection among privately evaluated model variants, yet neither the number of variants nor their dependence is public. We ask how many hidden variants a published margin can support while retaining statistical evidence of a provider's advantage over a fixed comparator. For a fixed candidate family under a Gaussian margin model, we derive a sensitivity curve that reports this maximum count as a function of a lower bound on within-family correlation. The relevant correlation must match the score used for ranking and the sampling model: in a controlled family, pooled item correlation is 0.90, whereas composite-score correlation is 0.46 under item resampling and 0.92 when MMLU subjects are resampled. An item-based audit of 394 adjacent-rank claims on the Open LLM Leaderboard finds that 391 lack statistical support even before accounting for selection. Among claims that pass the uncorrected test, certification can depend on assumptions about the hidden family's correlation. The resulting curves make these assumptions explicit without estimating the unobserved search size.
发表机构
- Texas A&M University(得克萨斯农工大学)
- Mayo Clinic(梅奥诊所)
机构由 AI 辅助整理,请以论文原文为准。