物理人工智能基准测试冗余的统计审计
A Statistical Audit of Physical AI Benchmark Redundancy
浏览论文内容
中文总结 AI 辅助
该研究通过构建51个模型在12个物理AI基准测试的矩阵,量化了基准测试的冗余性,提出贪婪选择基准测试的方法,得到4个基准测试子集保留78.5%效用并拟合排名,该方法适用于具足够重叠的基准测试。
中文摘要 AI 辅助
物理人工智能(Physical AI)模型在不同模型报告使用的基准测试套件上进行评估,导致模型-基准测试矩阵稀疏,且基准测试之间的关系未被测量。我们从包含51个基准测试和152个模型的注册表中,按报告密度选择12个物理AI基准测试,构建了包含51个模型的矩阵,结合模型卡片、基准测试论文的分数以及我们在每个基准测试官方协议下的自身评估运行结果。我们测量了基准测试共享的信息量,并提供了冗余的定量证据。冗余会影响报告的排名:将两对替代基准测试合并为单列后,在等权重平均下,51个模型中有22个的排名移动了3个或更多名次。随后,我们在结合分数离散度与未被已选集合解释的方差的效用指标下,贪婪选择基准测试,得到的4个基准测试子集保留了全部12个基准测试效用的78.5%,并在该子集上拟合了Bradley-Terry排名。该过程仅需要具有足够重叠的基准测试级分数,且并非特定于物理AI。
英文摘要
Physical AI models are evaluated on suites of benchmarks that differ across model reports, leaving the model-by-benchmark matrix sparse and the relationship between benchmarks unmeasured. We construct a matrix of 51 models on 12 physical AI benchmarks, selected from a registry of 51 benchmarks and 152 models by reporting density, combining scores from model cards and benchmark papers with our own evaluation runs under each benchmark's official protocol. We measure how much information the benchmarks share and show quantitative evidence of Redundancy. Redundancy affects reported rankings: collapsing the two substitute pairs into single columns moves 22 of 51 models by three or more places under an equally weighted average. We then select benchmarks greedily under a utility combining score dispersion with variance not explained by the already-selected set, and obtain a four-benchmark subset retaining 78.5\% of the utility of all 12, on which we fit a Bradley--Terry ranking. The procedure requires only benchmark-level scores with sufficient overlap and is not specific to physical AI.
发表机构
- Metric
机构由 AI 辅助整理,请以论文原文为准。