RECLAIM:智能体能否复现机器学习论文的声明?
RECLAIM: Can Agents Reproduce the Claims of Machine Learning Papers?
浏览论文内容
中文总结 AI 辅助
RECLAIM是一个由100篇NeurIPS 2025论文构成的基准,用于评估AI智能体复现论文结果的能力;实验表明,即使在最佳情况下,智能体在三个难度层级上的复现成功率也仅为41%、27%和15%。
中文摘要 AI 辅助
复现一篇机器学习论文涉及大多数研究步骤,从安装软件、调试到运行实验,这些工作越来越多地由AI智能体来完成。我们介绍了RECLAIM,一个包含100篇NeurIPS 2025论文的基准,该基准可每年从新会议中重建。对于每篇论文,我们预先确定要复现的结果、成功复现的标准以及GPU小时预算。智能体必须利用论文及其作者发布的任何内容来复现该结果。作者发布的内容决定了难度层级。Run层级发布包含代码、数据和权重;Retrain层级发布缺少权重,因此智能体需要训练模型;Reimplement层级发布缺少代码,因此智能体需要编写代码。一个独立的语言模型根据运行日志和输出(而非智能体的报告)进行评分。我们对每篇论文运行四个智能体各一次;每个层级中表现最好的智能体在Run层级仅复现了41%的论文,在Retrain层级为27%,在Reimplement层级为15%,而所有智能体在Reimplement层级表现最差。失败的尝试平均使用了29%的预算,因此大多数智能体在预算剩余时停止。最常见的智能体错误是在不将任何部分与论文中的数字进行核对的情况下编写方法,这在400次运行中出现了63次。
英文摘要
Reproducing a machine learning paper involves most research steps, from installing software and debugging to running experiments, work that AI agents increasingly do. We introduce RECLAIM, a benchmark of 100 NeurIPS 2025 papers that can be rebuilt yearly from new conferences. For each paper we fix in advance the result to reproduce, what counts as a successful reproduction, and a GPU-hour budget. An agent must reproduce that result using the paper and whatever its authors released. What the authors released decides the difficulty tier. Run-tier releases include code, data, and weights; Retrain-tier releases lack weights, so the agent trains the model; Reimplement-tier releases lack code, so the agent writes it. A separate language model grades runs from logs and outputs rather than agents' reports. We run four agents once per paper; the best agent in each tier reproduces only 41% of Run-tier papers, 27% at Retrain, and 15% at Reimplement, where every agent does worst. Failed attempts use on average 29% of their budget, so most stop with budget left. The most common agent error is writing the method without checking any part against the paper's numbers, in 63 of 400 runs.
发表机构
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
- National Center for Supercomputing Applications(国家超级计算应用中心)
机构由 AI 辅助整理,请以论文原文为准。