arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05179cs.CYcs.AI

自主研究智能体:AI科学家与验证差距的综述

Autonomous Research Agents: A Survey of AI Scientists and the Verification Gap

Tianyu Ding, Aditya Nannapaneni, Bingfan Liu, Ling Zhang

AI总结:

本综述聚焦AI/ML研究中AI科学家系统的主张验证差距,分析125项研究后纳入35项,发现多数系统发布代码但缺乏验证产物,贡献了语料库等成果,指出核心瓶颈转向主张的可验证性。

AI中文摘要:

大语言模型(LLM)智能体正越来越多地被应用于科学研究的全流程:创意构思、文献检索、实验设计与执行、结果分析、手稿撰写及同行评审。端到端的AI科学家系统如今已能产出类似论文的手稿,但其主张的验证难度往往高于代码运行的难度。本综述聚焦计算AI/ML研究中的这一差距,该领域的代码、基准、实验及成果记录最为显性。我们筛选了125项候选研究,纳入35项并对其中26项进行全文编码,包含24个可运行系统和2篇研究或立场论文。我们从7个审计维度进行编码:生命周期阶段、自主程度、评估方法、已发布产物、人在回路节点、新颖性验证及结果选择披露。主要模式为:代码发布如今已较为普遍,但可复现性等级产物及主张验证产物仍远不常见。在24个可运行系统中,83%发布了代码,38%发布了种子或执行轨迹,38%报告了任何新颖性验证方法;在9个闭环L4系统中,7个为机械重复,1个由作者宣称但无外部验证,本语料库中无LLM时代系统符合我们的编码规则,能展示经外部验证的在回路预言机。我们贡献了编码语料库、按自主程度划分的生命周期图谱、可审计性差距分析以及面向评审者的报告清单。本综述指出,该领域的核心瓶颈已不再仅仅是智能体能否完成研究任务,而是评审者能否验证这些智能体产出的主张。

英文摘要:

Large language model (LLM) agents are increasingly used across the scientific research lifecycle: ideation, literature search, experiment design and execution, analysis, manuscript drafting, and review. End-to-end AI scientist systems can now produce paper-like manuscripts, but their claims are often harder to verify than their code is to run. This survey studies that gap in computational AI/ML research, where code, benchmarks, experiments, and write-ups are most visible. We screen 125 candidate works and include 35, with full-text coding of 26 entries: 24 runnable systems and two study or position papers. We code seven audit dimensions: lifecycle stage, autonomy level, evaluation method, released artifacts, human-in-the-loop points, novelty verification, and result-selection disclosure. The main pattern is that code release is now common, but reproducibility-grade and claim-verification artifacts remain much less common. In the 24 runnable systems, 83 percent release code, while 38 percent release seeds or execution traces and 38 percent report any novelty-verification method. Among nine closed-loop L4 systems, seven are mechanical reruns and one is author-claimed without an external check; no LLM-era system in the corpus demonstrates an externally validated in-loop oracle under our coding rule. We contribute a coded corpus, a lifecycle-by-autonomy map, an auditability-gap analysis, and a reviewer-facing reporting checklist. The survey argues that the field's central bottleneck is no longer only whether agents can complete research tasks, but whether reviewers can verify the claims those agents produce.

↑