发表机构
Yale University(耶鲁大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MoRSE是一种面向任务的多智能体系统,通过角色-子任务专家混合机制与分层组相对策略优化,在代码生成基准上提升了任务及逐步性能,且专业化增益可跨任务类别与领域泛化。
AI 中文摘要
基于大语言模型的多智能体系统近期在复杂、长时程任务中展现出强大潜力。然而,现有方法主要依赖粗粒度的提示级区分,未针对不同子任务进行参数适配,导致智能体间异质性不足,专用能力有限,成为复杂需求任务性能的瓶颈。为解决这一问题,我们提出面向任务的混合角色-子任务专家多智能体系统(MoRSE),该系统在任务结构和参数层面均通过(角色,子任务)条件专业化来区分智能体。为在任务结构层面明确智能体的职责,我们构建了面向任务的多智能体系统,将每个任务分解为具有依赖关系感知的子任务有向无环图,并为每个智能体分配特定的(角色,子任务),在协作智能体间引入任务级专业化。此外,为应对不同角色和子任务的参数适配需求,我们提出了动态(角色,子任务)LoRA专家混合模块,该模块配备基于原型的子任务语义路由器,以高性价比的方式在共享大语言模型基座上增强智能体的参数级专业化。随后,为在稀疏任务奖励下稳定协同优化专家与路由器,我们进一步提出分层组相对策略优化方法,该方法包含两层信用分配机制,可将专家更新与路由决策引入的跨路由方差隔离开来,从而将专家质量与路由质量解耦。在三个主干模型上的代码生成基准测试实验验证了我们方法的有效性,该方法在整体任务性能和逐步性能上均有提升,且训练得到的专业化带来的增益可泛化到未见过的任务类别和领域。
英文摘要
Large language model-based multi-agent systems have recently shown strong potential for complex, long-horizon tasks. However, existing methods mainly rely on coarse prompt-level differentiation without parameter adaptation for diverse subtasks, resulting in insufficient inter-agent heterogeneity and limited specialized capability that bottleneck performance on tasks with complex requirements. To address this, we introduce a Task-Oriented Multi-Agent System with Mixture of Role-Subtask Experts (MoRSE) that distinguishes agents with (role, subtask)-conditional specialization at both the task structure and parameter levels. To make agents' responsibility explicit at the task structure level, we formulate a task-oriented multi-agent system that decomposes each task into a dependency-aware Directed Acyclic Graph of subtasks and assigns each agent a specific (role, subtask), introducing task-level specialization across collaborating agents. Additionally, to address the diverse role and subtask parameter adaptation demands, we propose a dynamic Mixture of (role, subtask) LoRA Experts module with a prototype-based semantic router for subtasks, augmenting agents with parameter-level specialization on a shared LLM substrate cost-effectively. Then, to co-optimize experts and router stably under sparse task rewards, we further propose a hierarchical group-relative policy optimization with two-layer credit assignment that isolates expert updates from the cross-route variance introduced by routing decisions, disentangling expert quality from routing quality. Experiments on code-generation benchmarks across three backbones demonstrate the effectiveness of our approach, with improvements in both whole-task and step-wise performance, and the gains from trained specialization generalize across held-out task categories and domains.
Comments25 pages, 8 figures, 9 tables