arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

深度学术调查:用于学术调查自动化的有状态智能体闭环范式

Deep Academic Survey: Stateful Agentic Closed-Loop Paradigm for Academic Survey Automation

Zhikai Xu, Zhucun Xue, Teng Hu, Yabiao Wang, Yong Liu, Jiangning Zhang

arXiv 2608.18034首次发表:更新:

发表机构

Zhejiang University; Shanghai Jiao Tong University(浙江大学; 上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出DAS框架,通过状态智能体闭环实现学术调查自动化,在DAS-Bench基准测试中,其在引用质量等四维度得分及专家评估中均优于现有系统。

AI 中文摘要

学术调查在组织快速扩张的学术文献中发挥核心作用,但其构建需要广泛的论文分析、连贯的知识组织、细粒度的引用支持以及可靠的手稿组装。现有的深度研究和自动调查生成系统仅能处理该过程的部分环节,通常无法通过共享的、可修订的状态协调论文理解、文献组织、基于证据的起草以及手稿验证。我们提出DAS,一个用于生成面向出版物的学术调查的有状态智能体框架,其核心思想是将可复用的论文分析与特定主题的手稿构建分离。DAS基于DAS-2M,一个包含约200万篇论文的面向调查的表示的动态更新元数据湖。其智能体通过基于候选的分类规划、反向论文到章节路由以及分层主张和引用规划,维护明确的文献、组织、写作和最终化状态。语义审查仅重新激活受影响的写作状态以进行修复和重新评估,形成具有确定性验证的范围闭环。我们进一步引入DAS-Bench,一个包含30个主题的基准,以及DAS-Eval,它通过16个标准评估学术引用质量、分类综合、分层论述和手稿组装可靠性。在所有30个主题上评估的系统中,DAS在所有四个维度均达到最高平均得分,总分为4.34,而最强竞争对手为4.03;在匹配的21个主题的计算机科学子集上,排序保持一致。盲法专家评估进一步显示,在30个主题中,27个主题的专家更偏好DAS而非Naive RAG,在21个共享计算机科学主题中,19个主题的专家更偏好DAS而非AutoSurvey。项目页面可通过此https URL访问。

英文摘要

Academic surveys play a central role in organizing rapidly expanding scholarly literature, yet their construction requires extensive paper analysis, coherent knowledge organization, fine-grained citation support, and reliable manuscript assembly. Existing Deep Research and automated survey generation systems address parts of this process, but typically do not coordinate paper understanding, literature organization, evidence-grounded drafting, and manuscript validation through a shared, revisable state. We introduce DAS, a stateful agentic framework for generating publication-oriented academic surveys. Its key idea is to separate reusable paper analysis from topic-specific manuscript construction. DAS builds on DAS-2M, a dynamically updated metadata lake containing survey-oriented representations of approximately two million papers. Its agents maintain explicit literature, organization, writing, and finalization states through candidate-grounded taxonomy planning, reverse paper-to-section routing, and hierarchical claim and citation planning. Semantic review reactivates only the affected writing states for repair and reevaluation, forming a scoped closed loop with deterministic validation. We further introduce DAS-Bench, a 30-topic benchmark, together with DAS-Eval, which assesses scholarly citation quality, taxonomic synthesis, hierarchical discourse, and manuscript assembly reliability through 16 criteria. Among systems evaluated on all 30 topics, DAS achieves the highest average in all four dimensions, with an overall score of 4.34 compared with 4.03 for the strongest competitor, and the same ordering is preserved on the matched 21-topic CS subset. Blinded expert evaluation further prefers DAS to Naive RAG on 27 of 30 topics and to AutoSurvey on 19 of 21 shared CS topics. The project page is available at https://zhikaixu24.github.io/projects/DAS/.

CommentsProject page: https://zhikaixu24.github.io/projects/DAS/ | Code: https://github.com/ZhikaiXu24/DAS | Data: https://huggingface.co/datasets/ZhikaiXu24/DAS-2M

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑