arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

训练AI科学家以实现研究的可复现性

Training AI Scientists to Replicate Research

Damon Falck, Samer Sabri, Anja Surina, Thom Foster, Anya Sims, Sam Devlin, Dylan Rogers, Tantum Collins, Kaloyan Aleksiev, Louis Kirsch, Edward Hughes

arXiv 2608.13331首次发表:更新:

发表机构

Inherent(因赫伦特)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究开发了可扩展的论文复现任务空间Replica,后训练出270亿参数的AI科学家Faraday,其在复现任务上表现优于Claude Opus 4.8和GPT-5.5,为长期科学创新AI智能体奠定基础。

AI 中文摘要

论文的可复现性是科学知识的基石,它确保了现有结果的可靠性,并为后续实验提供基础。复现行为通常会揭示之前未明确说明的细节,因此需要类似假设驱动的探索,以开展开放式研究。在本研究中,我们开发了Replica,这是一个可扩展的论文复现任务空间。为提供奖励信号,我们引入了一种自动生成的、基于评分标准的评判器,其噪声低且与人类对复现质量的评估一致。我们对Faraday进行了后训练,这是一个拥有270亿参数的“AI科学家”智能体,它利用编码智能体作为工具,在保留的复现任务上的表现超过了Claude Opus 4.8和GPT-5.5。对单个 rollout 的定性分析显示,Faraday采用了更符合科学原则的方法。我们认为,我们的结果为开发无需复杂框架、能够开展长期科学创新的AI智能体奠定了基础。

英文摘要

The replicability of papers is a cornerstone of scientific knowledge, ensuring the reliability of existing results and providing a base for further experiments. The act of replication typically illuminates details that were previously underspecified, and thus requires similar hypothesis-driven exploration to open-ended research. In this work, we develop Replica, a scalable task space for paper replication. To provide reward signal, we introduce an auto-generated rubric-based judge that has low noise and agrees with human assessment of replication quality. We post-train Faraday, a 27B-parameter "AI Scientist" agent that leverages coding agents as tools, surpassing the performance of Claude Opus 4.8 and GPT-5.5 on held-out replication tasks. Qualitative analysis of individual rollouts reveals that Faraday adopts a more scientifically-principled approach. We believe that our results provide a stepping stone towards AI agents capable of long-horizon scientific innovation without requiring complex harnesses.

Comments47 pages, 12 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑