发表机构
Inherent(因赫伦特)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究开发了可扩展的论文复现任务空间Replica,后训练出270亿参数的AI科学家Faraday,其在复现任务上表现优于Claude Opus 4.8和GPT-5.5,为长期科学创新AI智能体奠定基础。
AI 中文摘要
论文的可复现性是科学知识的基石,它确保了现有结果的可靠性,并为后续实验提供基础。复现行为通常会揭示之前未明确说明的细节,因此需要类似假设驱动的探索,以开展开放式研究。在本研究中,我们开发了Replica,这是一个可扩展的论文复现任务空间。为提供奖励信号,我们引入了一种自动生成的、基于评分标准的评判器,其噪声低且与人类对复现质量的评估一致。我们对Faraday进行了后训练,这是一个拥有270亿参数的“AI科学家”智能体,它利用编码智能体作为工具,在保留的复现任务上的表现超过了Claude Opus 4.8和GPT-5.5。对单个 rollout 的定性分析显示,Faraday采用了更符合科学原则的方法。我们认为,我们的结果为开发无需复杂框架、能够开展长期科学创新的AI智能体奠定了基础。
英文摘要
The replicability of papers is a cornerstone of scientific knowledge, ensuring the reliability of existing results and providing a base for further experiments. The act of replication typically illuminates details that were previously underspecified, and thus requires similar hypothesis-driven exploration to open-ended research. In this work, we develop Replica, a scalable task space for paper replication. To provide reward signal, we introduce an auto-generated rubric-based judge that has low noise and agrees with human assessment of replication quality. We post-train Faraday, a 27B-parameter "AI Scientist" agent that leverages coding agents as tools, surpassing the performance of Claude Opus 4.8 and GPT-5.5 on held-out replication tasks. Qualitative analysis of individual rollouts reveals that Faraday adopts a more scientifically-principled approach. We believe that our results provide a stepping stone towards AI agents capable of long-horizon scientific innovation without requiring complex harnesses.
Comments47 pages, 12 figures