发表机构
University of Science and Technology of China; Tencent; Zhejiang University; National University of Singapore(中国科学技术大学; 腾讯; 浙江大学; 新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多智能体系统协作训练难题,提出MAS-OPD在线策略蒸馏方法,通过角色优势专业化和协调特权归因提供密集监督,在代码与数学基准上取得最优平均分,提升角色专业化与协作效果。
AI 中文摘要
多智能体系统(MAS)将任务分配给专门角色,在复杂任务上具有前景,但一种主流方法仅依赖推理时的编排。通用API成本高昂且难以定制,而带有角色提示的小模型很少能发展出稳定的角色能力或可靠的协作,因此联合后训练一个MAS至关重要。大多数尝试使用强化学习,其团队级奖励无法确定是哪个代理的哪一步导致了结果,而局部奖励需要针对每个任务重新设计。在线策略蒸馏(OPD)在学生对轨迹进行采样时提供令牌级别的教师监督,这是一种更密集的信号,无需局部奖励,但尚未在MAS中相互依赖的代理上进行充分探索。出现两个困难:从行为属于哪个角色的判断中建立互补专业化,同时保留所有角色所需的知识;以及将跨代理协作信息转化为OPD可利用的监督。我们提出MAS-OPD,其中角色优势专业化将角色优势定义为目标角色条件和非目标角色条件下教师信号之间的差异,而协调的特权归因将交互冲突归因于其来源,并仅将其作为特权信息提供给教师。在代码和数学基准上的大量实验表明,MAS-OPD在两种学生规模下均获得最高平均分,并引导代理发展出更清晰的角色专业化和更有效的协作行为。
英文摘要
Multi-agent systems (MAS) split a task across specialized roles and are promising on complex tasks, yet a prevailing approach relies on inference-time orchestration alone. General-purpose APIs are costly and hard to customize, while small models with role prompts rarely develop stable role competence or reliable collaboration, so post-training a MAS jointly is central. Most attempts use reinforcement learning, whose team-level reward leaves undetermined which step of which agent brought about the outcome, while local rewards need redesigning per task. On-policy distillation (OPD) gives token-level teacher supervision on trajectories the student samples, a denser signal needing no local reward, yet is underexplored for the interdependent agents of a MAS. Two difficulties arise: building complementary specialization from a judgement of which role a behavior belongs to while preserving the knowledge all roles need, and turning cross-agent collaborative information into supervision OPD can exploit. We present MAS-OPD, where Role-Advantage Specialization defines the role advantage as the difference between the teacher signals under target and non-target role conditions, and Privileged Attribution for Coordination attributes an interaction conflict to its source and supplies it to the teacher alone as privileged information. Extensive experiments on code and mathematics benchmarks show that MAS-OPD attains the highest mean score at both student scales and leads the agents to develop clearer role specialization and more effective collaborative behavior.