arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ScientistTwo:以自主人工智能开拓人类知识前沿

ScientistTwo: Pioneering the Human Knowledge Frontier with Autonomous AI

Jaehyun Nam, Jinsung Yoon, Yanzhou Pan, Yubo Wang, Rui Meng, Parthasarathy Ranganathan, Tomas Pfister

arXiv 2609.19644首次发表:更新:

发表机构

Google Cloud AI Research; University of Waterloo(谷歌云AI研究院; 滑铁卢大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ScientistTwo是一个完全自主的多智能体框架,能够从初始问题出发,自主提出假设、设计并执行实验、通过模拟同行评审验证结果,生成专家级论文和可执行代码,其性能超越人类最先进模型。

AI 中文摘要

科学发现的核心在于识别现有知识的边界并探索未知领域。人工智能在科学领域的终极愿景是问题驱动的自主研究:在人类专家提出一个基础性挑战后,人工智能独立地在科学版图中导航,揭示理论与实证瓶颈,并系统地拓展知识前沿。本文介绍了ScientistTwo,一个旨在实现这一愿景的完全自主的多智能体框架。具体而言,ScientistTwo以初始问题为输入,建立最先进的基线,提出新颖假设,并协调专门智能体在无需人工干预的情况下编排端到端的发现周期。此外,该框架使用多样化的数据集和指标严格进行实验,通过自动化消融研究改进方法,并通过闭环模拟同行评审反驳引擎验证研究发现。为了以人类科学成就的最高标准评估ScientistTwo的能力,我们将其与ICLR、ICML和NeurIPS等顶级会议接受的论文进行基准测试。结果显示,ScientistTwo自主生成专家级、可发表的论文以及完全验证、可执行的代码库。其解决方案持续优于人类最先进的模型,并在自动化AI评审代理下获得比人类撰写的论文更高的平均评审评分。这些结果表明,ScientistTwo不仅仅是一个辅助工具,而是一个能够推动人类发现前沿的自主科学先驱。项目网站:此https URL。

英文摘要

Scientific discovery is defined by the ability to identify the boundaries of existing knowledge and venture into unexplored territory. The ultimate vision for AI in science is problem-driven autonomous research: given a fundamental challenge by a human expert, the AI independently navigates the scientific landscape, uncovers theoretical and empirical bottlenecks, and systematically expands the frontier of knowledge. In this paper, we introduce ScientistTwo, a fully autonomous multi-agent framework designed to realize this vision. Specifically, ScientistTwo takes an initial problem as input, establishes state-of-the-art baselines, formulates novel hypotheses, and coordinates specialized agents to orchestrate an end-to-end discovery cycle without human intervention. Moreover, the framework rigorously conducts experiments using diverse datasets and metrics, refines methodologies through automated ablation studies, and validates research findings via a closed-loop simulated peer-review rebuttal engine. To evaluate ScientistTwo's capabilities against the highest standards of human scientific achievement, we benchmark it across papers accepted at top-tier conferences such as ICLR, ICML, and NeurIPS. As a result, ScientistTwo autonomously generates expert-level, publishable papers and fully verified, executable codebases. Its solutions consistently outperform human state-of-the-art models, and achieve higher average review ratings than human-authored papers under automated AI review agents. These results show that ScientistTwo is not merely an assistive tool but an autonomous scientific pioneer capable of pushing the frontiers of human discovery. Project website: https://scientist-two.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑