arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

关于AI时间范围的估计与有效性——对METR图的统计学分析

On the estimation and validity of AI time horizons---a statistical look at the METR plot

Drew T. Nguyen, William Fithian

arXiv 2610.12466首次发表:更新:

发表机构

UC Berkeley(加州大学伯克利分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过样条函数和项目反应理论重新计算METR的50%AI时间范围,优化了估计方法并提供诊断图,为AI能力的可解释评估提供了更可靠的工具。

AI 中文摘要

METR的50%时间范围衡量AI以50%概率解决软件任务所需的人类完成时间,可将AI能力转化为可解释的单位。在228项任务和26种AI上,我们使用样条函数和项目反应理论重新计算时间范围,以放宽“任务的AI难度与人类时间的对数呈线性关系”的假设。我们拟合的样条函数可解释为将人类时间转换为AI难度的函数,其在2至30分钟区间内几乎平坦,其余区间接近线性。因此,尽管时间范围的乘数均为10倍,但从3分钟跃升至30分钟比从30分钟跃升至5小时容易得多。总体而言,我们提供了在交叉验证的适当评分规则集下表现更优的时间范围点估计值,以及用于评估时间范围构念效度的诊断图。我们建议,在提出新的基于时间范围的基准或现有基准纳入更长任务时,应结合诊断图解读时间范围。

英文摘要

METR's 50\% time horizon measures the human completion time of software tasks that an AI solves with 50\% probability, allowing AI capabilities to be expressed in interpretable units. On 228 tasks and 26 AIs, we recompute the time horizons using splines and item-response theory to relax the assumption that the AI difficulty of a task depends linearly on the log of human time. Our fitted spline can be interpreted as a function that \emph{converts} human time to AI difficulty; it is nearly flat in a region from 2--30 min but close to linear elsewhere. Hence, a time-horizon jump from 3 min to 30 min is much easier than one from 30 min to 5 hours despite the same multiplier of $10 \times$. Overall, we contribute time-horizon point estimates that perform better under a cross-validated suite of proper scoring rules, as well as diagnostic plots for assessing time horizons' construct validity. We suggest that time horizons be interpreted together with the diagnostic plots, especially as new time-horizon-based benchmarks are proposed or existing ones grow to include longer tasks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑