发表机构
Princeton University; Cornflower Labs; UK AI Security Institute; University of Toronto; Georgetown University (CSET); Johns Hopkins University; Golden Gate Institute for AI; Stanford University(普林斯顿大学; 矢车菊实验室; 英国人工智能安全研究所; 多伦多大学; 乔治城大学(CSET); 约翰斯·霍普金斯大学; 金门人工智能研究院; 斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究通过影子评估发现,前沿AI智能体可完成AI研究的工程工作,但在回答开放式研究问题时存在多种失败模式,无法取得实质性进展。
AI 中文摘要
AI爆炸式进步的预测依赖于AI智能体实现AI研究的自动化,但关于智能体能否开展开放式AI研究的证据不足。当前评估要么测试智能体在狭窄、可验证的任务上(排除了开放式研究),要么将AI生成的论文提交给盲审同行评审(该方式过度紧张、随机且评审质量差)。我们引入第三种方法来衡量AI研发自动化的进展:智能体承担一篇高质量未发表论文的核心开放式研究问题,由该论文的原作者对其输出进行评分,我们将这些称为影子评估。我们对两篇未发表的NeurIPS 2026投稿论文进行了影子评估,为前沿智能体提供了6天时间和数千美元的计算资源。智能体在无人帮助的情况下完成了所有工程工作,但在回答研究问题方面无法取得实质性进展,因此两篇论文都被原作者明确拒绝。我们确定了五种反复出现的失败模式:对可发表研究的标准判断不佳、对研究设计缺陷的回应缺乏创造性、从死胡同进行回溯效果不佳、资源意识差以及指令漂移。使用第二个模型和支架进行的稳健性检查重现了这些失败。我们发布了专家评审、调查回复、智能体仓库和日志。我们的结果提供了早期证据:当今的智能体可以完成AI研究的工程部分,但在研究生命周期的关键部分存在困难。
英文摘要
Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit AI-generated papers to blind peer review, which is overstretched, stochastic, and suffers from poor review quality. We introduce a third way to measure progress towards AI R\&D automation. An agent takes on the central, open-ended research question of a high-quality unpublished paper, and the paper's original authors grade its output. We call these shadow evaluations. We ran shadow evaluations on two unpublished NeurIPS 2026 submissions, giving frontier agents six days and thousands of dollars of compute. The agents completed all of the engineering without human help, yet could not make substantial progress towards answering the research questions. As a result, both papers were unambiguously rejected by the authors. We identify five recurring failure modes: poor judgment about the bar for publishable research, uncreative responses to shortcomings in the research design, ineffective backtracking from dead ends, poor resource awareness, and instruction drift. A robustness check with a second model and scaffold reproduced these failures. We release the expert reviews, survey responses, agent repositories, and logs. Our results provide early evidence that today's agents can do the engineering of AI research, but struggle with critical parts of the research lifecycle.