arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.09372cs.CL

MMLU 究竟测量了什么?聚合基准分数中难度结构的心理测量学审计

What Does MMLU Actually Measure? A Psychometric Audit of Difficulty Structure in Aggregate Benchmark Scores

  • Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

Dana Paquin, Riddhiman Jain

AI总结:

本文通过心理测量学审计证明 MMLU 聚合分数主要衡量事实检索而非推理能力,提出确定性框架并建议分项报告。

AI中文摘要:

尽管 MMLU 被广泛用作校准通用人工智能能力的基准,我们在心理测量学上证明其聚合分数主要评估的是模型的事实检索能力,而非推理能力。我们使用项目反应理论,对 1,000 个开放权重语言模型在 14,042 个 MMLU 测试项目上的项目难度进行校准,结果表明通过单一测试来评估这两种能力本身就存在缺陷。随后,难度被回归到一个确定性的、可文本提取的结构复杂度框架上。应用联合 Wald 检验并使用按学科聚类的协方差,我们证明 MMLU 混淆了根本可分离的构念。从结构复杂度到难度的映射在基准的 STEM 与非 STEM 分区之间并不具有不变性。这一发现具有实际后果。聚合排行榜排名与非 STEM 准确率的关联比与 STEM 准确率的关联更紧密,因此,为推理密集型部署而选择聚合分数上的前 50 名模型,会错失约 22% 的适合 STEM 任务的选项。此外,当在响应模型内部原生控制多项选择的猜测下限时,我们发现能力更高的模型在推理深度增加时性能下降得更陡峭。因此,MMLU 聚合分数对检索能力和推理稳定性权重不均,无意中偏向于为检索优化的模型。我们发布我们的确定性框架作为可复现的审计工具,并建议采用分项报告。

英文摘要:

Although MMLU is widely adopted as a benchmark for calibrating general AI capabilities, we psychometrically demonstrate that its aggregate score primarily evaluates a model's factual retrieval capacity rather than its reasoning ability. By calibrating item difficulty for 1,000 open-weights language models over 14,042 MMLU test items using Item Response Theory, we show that evaluating both abilities via a single test is inherently flawed. Difficulty is then regressed on a deterministic, text-extractable framework of structural complexity. Applying a joint Wald test with subject-clustered covariances demonstrates that the MMLU conflates fundamentally separable constructs. The mapping from structural complexity to difficulty is not invariant across the benchmark's STEM and non-STEM partitions. This finding has practical consequences. Aggregate leaderboard ranks track non-STEM accuracy more closely than STEM accuracy, so selecting a Top-50 model on the aggregate for a reasoning-intensive deployment displaces roughly 22% of the STEM-appropriate choices. Furthermore, when controlling for the multiple-choice guessing floor natively inside the response model, we find that higher-ability models continue to degrade more steeply under increased reasoning depth. The MMLU aggregate therefore weights retrieval capacity and reasoning stability unequally, inadvertently favoring models optimized for retrieval. We release our deterministic framework as a reproducible auditing instrument and recommend disaggregated reporting.

补充信息

↑