arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI智能体能否进行开放式科学发现?来自Station的证据

Can AI Agents Make Open-Ended Scientific Discovery? Evidence from Station

Wenyu Du, Stephen Chung

arXiv 2610.08927首次发表:更新:

发表机构

DualverseAI; University of Hong Kong; University of Cambridge(DualverseAI; 香港大学; 剑桥大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究在Station开放世界环境中验证AI智能体能否自主进行开放式科学发现,通过监督者机制和周期性元反思,使智能体平均重新发现62.7%的原始研究标准,显著优于现有基线,表明合适环境可支持智能体自主推进开放式科学发现。

AI 中文摘要

近期AI系统在给定明确指标的科学发现任务中取得了快速进展,但它们能否自主开展开放式科学发现仍不清楚。我们研究了AI在Station这一开放世界环境中的开放式任务能力,该环境中多个智能体模拟一个科学生态系统。为应对开放式任务特有的挑战,我们提出在Station中增加两种机制:监督者机制和周期性元反思,这些机制鼓励智能体在缺乏中间指标时仍持续探索。我们从ICLR近期发表的三篇口头报告中构建了开放式任务。我们向智能体提供每篇论文研究的主要问题,同时隐藏论文结果并禁用网络访问。随后我们衡量智能体重新发现了多少原始发现(按标准划分为独立标准)。我们发现Station平均重新发现了62.7%的标准,而Codex Multiagent-v2为15.4%,AI Scientist-v2为14.4%至20.6%。消融和行为分析表明,同时添加这两种机制可提高研究覆盖率和连续性。我们进一步在无参考论文的两个开放式任务上评估Station,发现智能体做出的一些发现与研究人员在知识截止日期后报告的发现高度吻合。综上,这些结果表明,合适的环境能使智能体在开放式科学发现中自主取得有意义的进展。

英文摘要

Recent AI systems have made rapid progress in scientific discovery when given well-defined metrics, but whether they can autonomously undertake open-ended scientific discovery remains unclear. We investigate AI's ability to tackle open-ended tasks in Station, an open-world environment in which multiple agents simulate a scientific ecosystem. To tackle challenges specific to open-ended tasks, we propose augmenting Station with two mechanisms: a Supervisor mechanism and periodic Meta Reflection, which encourage persistent exploration even when intermediate metrics are lacking. We construct open-ended tasks from three recent oral papers presented at ICLR. We give agents the main research question studied in each paper while withholding the paper's results and disabling web access. We then measure how many of the original findings-partitioned into individual criteria-agents rediscover. We find that Station rediscovers 62.7% of the criteria on average, compared with 15.4% for Codex Multiagent-v2 and 14.4-20.6% for AI Scientist-v2. Ablation and behavioral analyses indicate that adding the two mechanisms together improves research coverage and continuity. We further evaluate Station on two open-ended tasks without oracle papers and find that some of the discoveries made by the agents closely match discoveries reported by researchers after the knowledge cutoff date. Together, these results indicate that a suitable environment can enable agents to autonomously make meaningful progress in open-ended scientific discovery.

Commentspreprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑