arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TasteVal:衡量AI系统相对于人类专家的实验研究品味

TasteVal: Measuring the Experimental Research Taste of AI Systems Against Human Experts

Oliver Jaffe, Dane Sherburn

arXiv 2610.06824首次发表:更新:

发表机构

P-Zero Research(P-Zero Research)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

TasteVal通过计算效率衡量AI系统的实验研究品味,评估20个前沿模型,最佳模型超过人类专家基线,计算乘数每3个月翻倍,预测AI进展。

AI 中文摘要

我们推出了TasteVal,一个用于评估前沿模型实验研究品味的基准。我们将研究品味定义为挑选有趣问题、设计实验以及解读实验结果的能力。TasteVal衡量研究品味中的实验部分;在给定一个固定研究问题的情况下,我们衡量模型迭代设计实验并从中得出结论的能力。我们将实验研究品味操作化为计算效率;一个研究者若使用一半的串行实验计算量就达到与专家人类相同的分数,则其实验品味是后者的两倍。因此,实验品味充当实验计算量的乘数,使其成为预测AI进展的关键输入。TasteVal包含8个新颖、具有挑战性、开放式的任务,代表了前沿AI研发的典型特征。为了将品味与编码能力分离,被评估的模型扮演研究者的角色,迭代设计实验,而一个固定的编码智能体负责实现这些实验并报告结果。研究者持续执行,直到40个H100小时或120个挂钟小时的预算耗尽。我们招募了24位人类专家,每个任务至少2位,并取每个任务的最佳专家尝试作为专家基线。我们评估了2023年至2026年间发布的20个模型。表现最佳的模型Opus 5.5超过了我们的专家基线,计算乘数为2.3倍(95%置信区间1.15-4.37),其每次运行的平均成本约为我们基线人员的1/30。在TasteVal上,自2025年12月以来,前沿模型的计算乘数大约每3.0个月翻一番(95%置信区间1.7-5.0),而2023年至2025年12月期间为每14个月翻一番。以最终归一化性能衡量,前沿模型没有趋势断裂,每14.6个月翻一番。为保持TasteVal不受污染,我们不发布这些任务。

英文摘要

We introduce TasteVal, a benchmark to evaluate the experimental research taste of frontier models. We define research taste as the ability to pick interesting problems to solve, design experiments, and interpret experimental results. TasteVal measures the experimental component of research taste; given a fixed research problem, we measure how well a model iteratively designs experiments and draws conclusions from their outcomes. We operationalize experimental research taste as compute efficiency; a Researcher who reaches the same score as an expert human using half the serial experimental compute has twice the experimental taste. Experimental taste thus acts as a multiplier on experimental compute, making it a key input to forecasts of AI progress. TasteVal consists of 8 novel, challenging, open-ended tasks representative of frontier AI R&D. To isolate taste from coding ability, the model under evaluation acts as a Researcher that iteratively designs experiments while a fixed Coder agent implements them and reports their results. The Researcher executes until either the 40 H100 hour or 120 wall-clock hour budgets are exhausted. We recruit 24 human experts, at least 2 per task, and take the best expert attempt per task as the expert baseline. We evaluate 20 models released between 2023 and 2026. The best-performing model, Opus 5.5, exceeds our expert baseline, with a compute multiplier of 2.3x (95% CI 1.15-4.37), at roughly 1/30 of our baseliners' average per-run cost. On TasteVal, the compute multiplier of frontier models has doubled approximately every 3.0 months since December 2025 (95% CI 1.7-5.0), up from every 14 months between 2023 and December 2025. Measured by final normalized performance, frontier models show no trend break, doubling every 14.6 months. To keep TasteVal uncontaminated, we do not release the tasks.

Comments38 pages, 21 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑