arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DAGent: 面向深度研究智能体的评估-再扩展规划

DAGent: Evaluate-then-Grow Planning for Deep Research Agents

Hanwen Liu, Yuanfu Sun, Qiaoyu Tan

arXiv 2609.39154首次发表:更新:

发表机构

New York University; New York University Shanghai(纽约大学; 上海纽约大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DAGent提出评估-再扩展规划框架,通过增量图构建和拓扑强化学习提升深度研究智能体的准确性与效率,在多个基准上超越强基线。

AI 中文摘要

深度研究任务要求智能体在庞大的知识空间中导航,综合多个来源的证据,并随着发现的涌现调整其计划。基于有向无环图(DAG)的多智能体系统适合这一场景,因为它们支持并行执行,并将每个子任务隔离在聚焦的依赖上下文之中。然而,现有的基于DAG的智能体在执行前实例化任务级计划,并且仅在观察到失败或缺失证据后才修复图。这种“先计划后修补”的策略对于深度研究而言是脆弱的:系统在证据最薄弱时承诺最强,而后续的修订浪费了本不应规划的分支上的计算。我们提出DAGent,一个基于DAG的多智能体框架,采用“评估-再扩展”增量规划:一个编排器逐批扩展任务图,每次扩展都基于已完成节点的置信度和不确定性信号。一个分层上下文层默认传播紧凑的QueryDocs,同时保留完整的执行轨迹以供按需回忆。记录的DAG拓扑允许结构化的强化学习信号,而仅基于结果的配方无法定义这些信号;DAGRPO,一种GRPO的改编,将拓扑条件信用注入执行器回滚,并对编排器计划施加结构合规正则化。在BrowseComp-Plus、GAIA和xbench-DeepSearch上,DAGent在Qwen3-235B-A22B规模上超越最强的开源基线5.3 / 5.8 / 2.0个百分点,且领先优势在四个开源骨干上复制,并扩展到327K上下文下的GPT-5。在Qwen3-8B规模上,DAGRPO比同预算的仅基于结果的GRPO基线平均Pass@1提高3.0个百分点。一项同架构比较表明,证据条件规划在较低的单任务token、工具调用和步骤足迹下达到比其“先计划后修补”对应物更高的准确率。代码:此HTTPS URL

英文摘要

Deep research tasks require agents to navigate large knowledge spaces, synthesize evidence across many sources, and adapt their plans as findings emerge. Directed acyclic graph (DAG)-based multi-agent systems suit this setting because they support parallel execution and isolate each sub-task within a focused dependency context. Yet existing DAG-based agents instantiate a task-level plan before execution and repair the graph only after failures or missing evidence are observed. This Plan-then-Patch strategy is brittle for deep research: the system commits most strongly when its evidence is weakest, and later revisions waste computation on branches that should not have been planned. We propose DAGent, a DAG-based multi-agent framework with Evaluate-then-Grow incremental planning: an Orchestrator grows the task graph one batch at a time, conditioning each expansion on confidence and uncertainty signals from completed nodes. A hierarchical context layer propagates compact QueryDocs by default while preserving full execution traces for on-demand recall. The recorded DAG topology admits structural RL signals that outcome-only recipes cannot define; DAGRPO, a GRPO adaptation, injects topology-conditioned credit on Executor rollouts and a structural compliance regularization on Orchestrator plans. Across BrowseComp-Plus, GAIA, and xbench-DeepSearch, DAGent surpasses the strongest open-source baseline by 5.3 / 5.8 / 2.0 points at the Qwen3-235B-A22B scale, and the lead replicates across four open-source backbones and extends to GPT-5 at 327K context. At the Qwen3-8B scale, DAGRPO improves over a same-budget outcome-only GRPO baseline by 3.0 average Pass@1 points. A same-architecture comparison shows that evidence-conditioned planning reaches higher accuracy at lower per-task token, tool-call, and step footprints than its Plan-then-Patch counterpart. Code: https://github.com/hanwenliu6825/DAGent

CommentsAccepted at NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑