智能体在自动研究(AutoResearch)中如何失败?针对100项真实前沿研究任务的端到端诊断评估
How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks
浏览论文内容
中文总结 AI 辅助
本研究推出AutoResearchEval评估平台,构建含45种失效模式的ARFT分类体系,发现当前智能体缺乏元认知循环是核心缺陷,公开发布相关资源以推动自主科学发现研究。
中文摘要 AI 辅助
长期以来,AI一直为科学研究提供辅助,但大型语言模型(LLM)和智能体框架的快速发展正在重塑这一格局;如今单个系统已能完成从初始假设到最终发表论文的全阶段研究,这一范式被称为自动研究(AutoResearch)。现有评估几乎未揭示这些智能体的运作方式或失效环节:任务范围狭窄,评估仅衡量性能而非过程,失效诊断缺乏系统性覆盖或人工制品层面的可见性。为解决这一缺口,我们推出AutoResearchEval,包含100项基于已发表前沿科学的任务,覆盖7个科学领域及完整研究生命周期,包括构思、检索、执行、分析、写作和评审环节。对8种框架-模型组合进行评估,共生成800条自动研究智能体轨迹,并附带过程级标注。我们将这些见解整理为自动研究失效分类体系(AutoResearch Failure Taxonomy,简称ARFT),这是一个包含45种基于实证的失效模式的框架。为实现可扩展的细因归因,我们采用经人工校准的智能体作为评判者的流程,检查完整轨迹和中间人工制品。失效模式汇聚为一个核心局限,即当前智能体缺乏元认知循环,元认知循环指的是将自身产出与所发现内容对照检查、在不成立时进行修正、并质疑自身采取路径是否合理的能力。相同模式在全部8种框架-模型组合中反复出现,包括测试的最强模型,这一缺陷定位于模型层面而非特定框架; orchestration层面的干预是否能弥补这一缺陷,是本研究未验证的开放性问题。我们公开发布AutoResearchEval和ARFT,以推动自主科学发现领域的持续研究与开发。
英文摘要
AI has long assisted scientific research, but the rapid advance of LLMs and agentic scaffolds is reshaping the landscape; a single system can now carry whole-stage research from an initial hypothesis all the way to final published paper, which is a paradigm now referred to as AutoResearch. Existing evaluations reveal little about how these agents operate or where they break down. Tasks are narrowly-scoped, evaluation measures performance but not process, and failure diagnoses lack systematic coverage or artifact-level visibility. To address this gap, we introduce AutoResearchEval, featuring 100 tasks grounded in published frontier science across 7 scientific domains and the full research lifecycle, including ideation, retrieval, execution, analysis, writing, and review. Evaluating 8 harness-model combinations yields 800 autoresearch agent trajectories, with process-level annotation. We organize these insights into AutoResearch Failure Taxonomy or ARFT, a framework of 45 empirically-grounded failure patterns. To enable scalable fine-grained attribution, we leverage a human-calibrated agent-as-a-judge pipeline to inspect complete trajectories and intermediate artifacts. Failure patterns converge on a single overarching limitation, namely that current agents lack a metacognitive loop, which entails the ability to check what they produced against what they found, revise when it does not hold up, and question whether the path they took was sound. The same patterns recur across all 8 harness-model combinations, including the strongest models tested, locating the deficit at the model level rather than in any particular scaffold; whether orchestration-level interventions can close it is an open question this work does not test. We publicly release AutoResearchEval and ARFT to facilitate continued research and development in autonomous scientific discovery.