AI 中文总结
研究量化语言模型基准测试中排名不确定性问题,通过汇总成对假设检验来实现,分析了知识评估基准MMLU的不确定性来源并展示如何修改假设检验,指出MMLU各主题排名变异性大,比较模型时应考虑。
AI 中文摘要
预训练模型通常在多任务排行榜上排名以评估其在不同任务中的有效性。最近引入了排名置信区间,通过汇总成对假设检验来量化这些排名中的不确定性。在这项工作中,我们分析了知识评估基准MMLU中的不确定性来源,并展示了如何修改假设检验以考虑其影响。我们证明,MMLU各主题间的排名变异性很大,在比较语言模型或识别最佳模型时应予以考虑。
英文摘要
Pretrained models are typically ranked on multi-task leaderboards to assess their effectiveness across diverse tasks. Rank confidence intervals were recently introduced as a method to quantify the uncertainty in these rankings by aggregating pairwise hypothesis tests. In this work, we analyze the sources of uncertainty in the knowledge evaluation benchmark MMLU and show how hypothesis tests can be modified to account for their effects. We demonstrate that ranking variability across MMLU subjects is substantial and should be considered when comparing LLMs or identifying the top-performing models.