arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MAGIC:基于密集奖励强化学习增量构建的混合粒度智能体图

MAGIC: Mixed-Granularity Agent Graphs via Incremental Construction with Dense-Reward Reinforcement Learning

Kairui Yang, Ziheng Yi, Xunkai Li, Minghao An, Zhanke Liu, Zekai Chen, Rong-Hua Li

arXiv 2609.26667首次发表:更新:

AI 中文总结

针对多智能体系统协作拓扑生成忽视子任务粒度差异的问题,提出MAGIC框架,通过密集奖励强化学习增量构建混合粒度智能体图,在八个基准上超越现有方法并提升推理效率。

AI 中文摘要

协作拓扑结构决定了基于大语言模型的多智能体系统的性能和执行成本。由于任务在复杂度和所需能力上各不相同,近期方法会生成特定于任务的协作图,以指定智能体的参与和信息流。然而,代表性的拓扑生成器在整个组织中使用单个智能体或预定义组,忽视了不同子任务间协作需求的差异。我们的关键洞察是,为每个功能角色局部选择粒度,在单一组织内结合细粒度控制与可复用的协作模式。学习此类组织需要在一个组合构建空间中进行探索,而最终答案奖励提供的中间反馈有限。因此,我们提出MAGIC,一个用于混合粒度图生成的密集奖励强化学习框架。具体而言,MAGIC通过顺序选择功能角色、将其实例化为单个智能体或可复用组,并连接到现有单元,来构建混合粒度智能体图。我们直接利用当前策略下采样轨迹的回报来优化构建策略,并使用基于势能的奖励塑形,从基于探针的效用和结构信号中提供中间反馈,同时保留累积任务奖励。MAGIC在八个基准上优于最先进的基线,并在效率研究中展现出强大的推理效率。

英文摘要

Collaboration topology shapes both the performance and execution cost of LLM-based multi-agent systems. Because tasks differ in complexity and required capabilities, recent approaches generate task-specific collaboration graphs that specify agent participation and information flow. However, representative topology generators use either individual agents or predefined groups throughout an organization, overlooking differing collaboration needs across subtasks. Our key insight is to select granularity locally for each functional role, combining fine-grained control with reusable collaboration patterns within one organization. Learning such organizations requires exploring a combinatorial construction space with limited intermediate feedback from final-answer rewards. Therefore, we propose MAGIC, a dense-reward reinforcement learning framework for mixed-granularity graph generation. Specifically, MAGIC constructs a mixed-granularity agent graph by sequentially selecting a functional role, instantiating it as a single agent or reusable group, and connecting it to existing units. We directly optimize the construction policy using returns from trajectories sampled under the current policy and use potential-based reward shaping to provide intermediate feedback from probe-based utility and structural signals while preserving the cumulative task reward. MAGIC outperforms state-of-the-art baselines across eight benchmarks and demonstrates strong inference efficiency in our efficiency study.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑