arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29420cs.LGcs.CYcs.SE

一种能力还是多种?测试前沿AI评估的经济有效性

One Capability or Many? Structural and Predictive Tests of Benchmark Validity Disagree About Economic Benchmarks for Frontier AI

Louis Yiven Zhu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究测试前沿AI评估的经济基准是否对应独立能力,发现其未形成独特因子,仅为日期驱动的通用能力因子提供增量信息,同期模型差距需经日期调整后才反映能力差异。

中文摘要 AI 辅助

前沿模型排行榜目前基于经济基准、模型执行从软件工程到银行业务流程等专业任务的测试来对系统进行排名,这些排名影响着企业的采购决策、监管机构的审查重点,以及人们对工作将如何变化的预期。这类基准衡量的是一种区别于通用应试能力的独特能力,还是仅反映了模型性能提升过程中所有基准都会沿之上升的单一维度,这是一个尚未得到研究的构念效度问题。我们在一个哈希固定的排行榜快照上对该问题进行了测试,该快照包含12个基准(其中4个为经济基准)下的421种模型配置,将基准视为项目、模型视为被试,采用潜在变量模型,预先设定4个假设及其阈值。单一因子解释了74.5%的共同方差,且与模型发布日期相关(R²=0.505),因此能力的主导维度在很大程度上是时间趋势;在控制规模的前提下,若剔除日期因素,计算量的影响微乎其微。剔除日期趋势会使该占比降低14.9个百分点,若每个基础模型仅保留一行数据则降低24.1个百分点。根据预先设定的维度规则,经济基准未形成独立因子;但采用留一基准法测试(在每个折内重新估计因子)显示,多因子表示比单一通用指数能更好地预测留出的经济得分(合并Delta-MSE为0.037,95%自助抽样区间为[0.019, 0.055])。因此,经济基准为一个主要由日期驱动的通用因子提供了增量预测信息,证据不支持将其视为独特的潜在能力。排行榜仍是整体进展的可靠指南,但相隔数月发布的模型间的大部分差距源于时间,因此同期模型间的较小差距应先进行日期调整,再解读为能力差异。

英文摘要

Frontier-model leaderboards now rank systems on economic benchmarks, and those rankings inform what organisations buy and what regulators scrutinise. Whether such benchmarks measure a capability distinct from general test-taking is a question of construct validity that a structural test and a predictive test can answer in opposite ways. We show that they do on a hash-pinned snapshot of a frontier leaderboard with 421 model configurations across twelve benchmarks, four of them economic, of which 103 configurations carry all three sparsely scored economic benchmarks and 96 carry all twelve; four hypotheses and their thresholds were fixed before analysis, and every deviation from the plan is reported. The first factor of a three-factor extraction carries 74.5% of common variance and tracks release date (R^2 = 0.505), and date adjustment lowers its share by 14.9 points. Under the dimensionality rule fixed in advance the economic benchmarks form no factor of their own. A leave-one-benchmark-out test with factors re-estimated inside every fold nevertheless finds that a multi-factor representation predicts held-out economic scores better than a single general index (pooled Delta-MSE 0.037, 95% bootstrap interval [0.019, 0.055]) under the linear learners that fit best, an advantage that reverses for tree learners. Under the linear learners the same representation also predicts the eight other benchmarks better, so the battery carries predictive structure that one index misses and the economic benchmarks share it without forming a distinct factor. Construct validity should therefore be assessed by predictive tests alongside structural ones. We give a two-test protocol for benchmark builders and release the pinned data, the analysis plan and the code.

发表机构

  • University of Oxford(牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑