arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

IdeaScientist:编排智能体实现有依据的科学构思

IdeaScientist: Orchestrating Agents for Grounded Scientific Ideation

Jiarui Liu, Renjie Tao, Yiwei Liao, Chuanyang Jin, Kai Sun, Xiao Yang, Xinyuan Zhang, Xilun Chen, Zhuangqun Huang, Lechen Zhang, Yongjin Yang, Yinghui He, Weihao Xuan, Rakesh Wanga, Anuj Kumar, Mona T. Diab, Wen-tau Yih, Xin Luna Dong

arXiv 2610.04074首次发表:更新:

发表机构

Carnegie Mellon University; Meta; Meta Reality Labs; FAIR at Meta; University of Illinois Urbana-Champaign; University of Toronto; Princeton University; The University of Tokyo; RIKEN AIP(卡内基梅隆大学; Meta; Meta Reality Labs; Meta FAIR; 伊利诺伊大学厄巴纳-香槟分校; 多伦多大学; 普林斯顿大学; 东京大学; 日本理化学研究所人工智能中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

IdeaScientist 通过强化学习训练三个角色(发现空白、创新、报告撰写),并利用 Svalbard Idea Vault 语料库,在时间控制评估中显著超越开源基线,甚至超过闭源模型,实现有依据的科学构思。

AI 中文摘要

尽管自动化科学研究取得了快速进展,但生成有前景且依据充分的研究解决方案仍是一个核心挑战。我们将研究构思(research ideation)独立为一个任务,并基于以下直觉构建解决方案:一个领域的挑战往往可以通过另一个领域中解决类似挑战的机制来应对。据此,我们提出了 IdeaScientist,它将构思过程分解为发现空白(gap finding)、创新(innovation)和报告撰写(report writing),并使用强化学习训练每个角色。这些角色识别相关工作中的局限性,从类似问题情境中汲取解决思路,并将这些思路发展成完整的研究提案。为了促进跨领域的见解发现,我们构建了 Svalbard Idea Vault,这是一个包含 277 万个分解研究想法的语料库,用于检索、训练和时间控制评估。我们的评估限制访问截止日期之前可用的文献,并评估所提出的方向与人类研究人员后来在 15K 篇论文中探索的方向的吻合程度。在 Qwen3.6-27B 上,IdeaScientist 比最强的开源自动研究基线高出 14.0%,主要得益于新颖性的提升。在这个 27B 开源骨干上,IdeaScientist 甚至比使用 Claude-4.8-Opus 的 Claude Code SDK 和使用 GPT-5.4 的 Codex SDK 高出最多 5.9%。

英文摘要

Despite rapid progress in automating scientific research, generating promising and well grounded research solutions remains a central challenge. We isolate research ideation as a standalone task and build our solution on the intuition that a challenge in one field can often be addressed by a mechanism that solved an analogous challenge in another. Accordingly, we introduce IdeaScientist, which decomposes ideation into gap finding, innovation, and report writing, and trains each role with reinforcement learning. These roles identify limitations in related work, draw solution intuitions from analogous problem settings, and develop those intuitions into complete research proposals. To facilitate discovery of insights across domains, we construct the Svalbard Idea Vault, a corpus of 2.77M decomposed research ideas for retrieval, training, and temporally controlled evaluation. Our evaluation restricts access to literature available before a cutoff date and assesses how closely proposed directions align with those later explored in 15K papers authored by human researchers. On Qwen3.6-27B, IdeaScientist outperforms the strongest open-source autoresearch baseline by 14.0%, driven mainly by gains in novelty. On this 27B open backbone, IdeaScientist even outperforms Claude Code SDK with Claude-4.8-Opus and Codex SDK with GPT-5.4, by up to 5.9%.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑