arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

前沿AI预测存在测量问题:对进展证据的审计

Frontier AI Forecasting Has a Measurement Problem: An Audit of Progress Evidence

Fabricio F Costa

arXiv 2608.14903首次发表:更新:

发表机构

AIx4All, LLC; HCLTech; Stanley Manne Children’s Research Institute; Ann & Robert H. Lurie Children’s Hospital of Chicago; Northwestern University Feinberg School of Medicine(AIx4All有限责任公司; HCLTech公司; 斯坦利·曼恩儿童研究所; 安与罗伯特·H·卢里芝加哥儿童医院; 西北大学费恩伯格医学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文审计前沿AI预测的测量问题,构建含多类元素的冻结记录,发现训练计算量、METR horizon等关键数据缺失,基准更迭存在断点,来源高度集中,提出需明确版本化测量系统的预测主张。

AI 中文摘要

前沿人工智能的定量预测常将过时目标与基准分数、训练计算量、发布时间或专家信念的趋势关联起来。本文在拟合另一条趋势前,审计公开测量记录是否支持这些关联。我构建了截至2026年8月12日、以事件为中心的冻结记录,包含62个选定系统、12个版本化基准、7个能力或影响标准、144个分级事件、27个来源记录和408个类型化关系。该记录是审计样本而非普查。仅7个系统同时具备估计训练计算量和METR 50%任务 horizon,27个闭源系统中有19个缺失训练计算量,包括2026年所有选定的闭源发布,而35个开放权重系统中无一个有METR horizon观测。基准更迭造成第二个断点:METR Time Horizon 1.0到1.1的7个系统链路,对数尺度斜率为1.206(95%置信区间1.021至1.390);而MMLU到MMLU-Pro的6个系统比较,在logit和probit链接下呈偏移状,在线性或对数链接下则不然。观测到的链路仅对接近25%的斜率偏离有约80%的功效。来源集中度高:71个实质性定量事件中的52个(73.2%)来自一个测量计划,76.1%为实验室发布。对56个方法学和经验来源的审查确定了16个补充测量方向,涵盖资源、推理预算、可靠性、智能体工作、安全、人类偏好、领域结果和预测回测,但无方向提供替代标量。结果并非前沿AI预测不可能,而是可辩护的过时预测是关于版本化测量系统的主张,该系统具有明确的连接、协议、链路和来源依赖,而非仅拟合曲线或日历日期。

英文摘要

Quantitative forecasts of frontier artificial intelligence often connect dated targets to trends in benchmark scores, training compute, release time, or expert belief. This paper audits whether the public measurement record supports those connections before another trend is fitted. I construct a frozen, event-centric record through 12 August 2026 with 62 selected systems, 12 versioned benchmarks, seven capability or impact criteria, 144 graded events, 27 source records, and 408 typed relations. The record is an audit sample, not a census. Only seven systems jointly observe estimated training compute and a METR 50 percent task horizon. Training compute is absent for 19 of 27 closed systems, including every selected closed release from 2026, while none of the 35 open-weight systems has a METR horizon observation. Benchmark succession creates a second break: a seven-system link from METR Time Horizon 1.0 to 1.1 has a log-scale slope of 1.206 (95 percent CI 1.021 to 1.390), whereas a six-system MMLU to MMLU-Pro comparison appears shift-like under logit and probit links but not under linear or logarithmic links. The observed bridges have about 80 percent power only for slope departures near 25 percent. Provenance is concentrated: 52 of 71 substantive quantitative events, or 73.2 percent, come from one measurement programme, and 76.1 percent are laboratory releases. A review of 56 methodological and empirical sources identifies 16 complementary measurement directions spanning resources, inference budgets, reliability, agentic work, safety, human preference, field outcomes, and forecast backtesting. No direction supplies a replacement scalar. The result is not that frontier AI forecasting is impossible, but that a defensible dated forecast is a claim about a versioned measurement system with explicit joins, protocols, links, and source dependence, not merely a fitted curve or calendar date.

Comments16 pages, 6 figures, 1 table. Evidence cutoff: 12 August 2026. Ancillary code and data included. Preprint; comments welcome

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑