arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DHCG:面向基于LLM的多智能体推理的分层协作图动态构建

DHCG: Dynamic Construction of Hierarchical Collaboration Graphs for LLM-Based Multi-Agent Reasoning

Jie Ren, Jiakang Yuan, Chenyu Huang, Hezeer Ma, Jiayuan Fan, Tao Chen

arXiv 2610.07835首次发表:更新:

发表机构

Fudan University; Shanghai Innovation Institute(复旦大学; 上海创新研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DHCG将多智能体推理建模为部分可观测马尔可夫决策过程,动态构建分层协作图,通过规划器、执行器和生成器协作,在代码生成等基准上平均性能领先,较单智能体提升13.06点。

AI 中文摘要

基于LLM的多智能体系统(MAS)在解决跨领域复杂问题方面已展现出强大的能力。近期,智能体系统的动态编排已成为一个重要的研究方向。然而,现有方法存在组合受限、依赖关系错位和规模不灵活等问题,限制了其在执行过程中适应推理需求的能力。为解决这些局限,我们将MAS设计重新构建为一个部分可观测马尔可夫决策过程,其中MAS的组合和规模均被动态确定。我们提出DHCG,一个新颖的框架,协调三个模块(规划器、执行器和生成器),基于查询和不断演化的执行反馈,从零开始逐步构建动态分层协作图。在每一步中,在反馈的引导下,规划器生成一组针对当前推理需求定制的独特且互补的角色,并选择性地将相关信息路由至每个角色。它还可以提前终止分层协作图,或在需要额外推理时逐步扩展该图。我们进一步引入行动感知偏好优化,以训练规划器在构建分层协作图时做出更有效的决策。我们在代码生成、数学推理和领域特定推理基准上系统评估了DHCG。DHCG在比较方法中取得了平均性能的最先进水平,比单智能体基线提高了13.06个百分点,并比静态和动态MAS基线高出2.77-8.02个百分点。额外实验进一步证明了其在不同规划器骨干和未见执行器模型上的泛化能力。

英文摘要

LLM-based multi-agent systems (MAS) have demonstrated strong capabilities in solving complex problems across diverse domains. Recently, the dynamic orchestration of agent systems has become an important research direction. However, existing methods suffer from limited composition, misaligned dependencies, and inflexible scale, restricting their ability to adapt to reasoning requirements during execution. To address these limitations, we reframe MAS design as a partially observable Markov decision process, in which both the composition and scale of the MAS are dynamically determined. We propose DHCG, a novel framework that coordinates three modules (Planner, Worker, and Generator) to progressively construct a dynamic hierarchical collaboration graph from scratch based on the query and evolving execution feedback. At each step, guided by feedback, the Planner generates a set of distinct and complementary roles tailored to the current reasoning needs and selectively routes relevant information to each role. It can also finalize the hierarchical collaboration graph early or progressively expand it when additional reasoning is required. We further introduce action-aware preference optimization to train the Planner to make more effective decisions when constructing hierarchical collaboration graphs. We systematically evaluate DHCG across code generation, mathematical reasoning, and domain-specific reasoning benchmarks. DHCG achieves state-of-the-art average performance among the compared methods, improving over the single-agent baseline by 13.06 points and outperforming both static and dynamic MAS baselines by 2.77-8.02 points. Additional experiments further demonstrate its generalization across different Planner backbones and unseen Worker models.

Comments9 pages, 4 figures, 4 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑