AI 中文总结
本文提出适用于共享评估的多语言模型比较统计方法,构建含模型固定效应、问题随机效应的随机效应模型,经模拟与真实数据验证,为多模型评估结果报告提供建议。
AI 中文摘要
针对同一评估任务下两个模型的比较,可通过配对t检验分析其标准误并对相关问题进行聚类校正,从而实现严谨的统计处理。不过,排行榜、消融研究及超参数搜索通常会同时比较K>2个模型。本文提出一种针对共享评估得分的单一随机效应模型,将模型作为固定效应、问题(或问题簇)作为随机效应。研究表明,通过经典方差分析或线性混合模型拟合该模型,可在K=2时恢复Miller的配对与聚类估计量,同时将结果扩展至任意K值及不平衡、聚类设计。通过模拟研究和实际数据应用验证了该模型,采用6个公开可用的语言模型,在14个学科簇的1497个共享MMLU-Pro问题上评分,结果显示,成对排序结论的有效性取决于是否恰当考虑问题层面的配对与多重比较。文末提供了报告多模型评估结果的具体建议。
英文摘要
A rigorous statistical treatment of two-model comparisons on the same evals can be achieved by paired t-tests, analyzing their standard errors and a clustering correction for correlated questions. Nevertheless, leaderboards, ablations studies, and hyperparameter sweeps, usually compare $K>2$ models simultaneously. In this paper, we present a single random-effects model for scores on shared evals. We leverage models as a fixed effects and questions (or question cluster) as random effects. We show that fitting it a classical ANOVA or a linear mixed model we can recover Miller's paired and clustered estimators for K=2, but we extend the results to any K and to unbalanced, clustered designs. We validate the described model in a simulation study and using a real-data application. We take six openly available language models scored on 1,497 shared MMLU-Pro questions across 14 subject clusters, and we show that pairwise ranking claims survive depending on properly accounting for both question-level pairing and multiple comparisons. By the end of the manuscript, we provide concrete recommendations for reporting multi-model eval results.