arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.08452cs.AI

SRPO:用于多智能体大语言模型的集合级相对策略优化

SRPO: Setwise Relative Policy Optimization for Multi-Agent Systems

Shengtian Yang, Ziyu Xiong, Yu Li, Yewen Li, Qingpeng Cai, Lei Feng

首次发表
浏览论文内容

中文总结 AI 辅助

SRPO将多智能体LLM中一次转换消耗的输出集视为一个动作,通过集合级相对策略优化统一分工与共同进化,在数学推理和多轮搜索中跨四种模型规模取得最优宏平均结果。

中文摘要 AI 辅助

多智能体大语言模型通过在共享环境中协调多个策略来解决复杂任务。然而,现有的强化学习方法通常单独优化每个响应或轨迹,即使多个输出共同导致一次状态转换。因此,更新单元与系统执行的动作不一致。为解决此问题,我们提出SRPO(集合级相对策略优化),将活动集(一次转换所消耗的最小输出集)视为一个多智能体动作。具体而言,SRPO将成员对数比率组合为一个基数归一化的集合比率,分配一个相对优势,并一次性裁剪该集合。这种表述将分工和联合共同进化统一为具有不同集合大小的动作。在数学推理和多轮搜索上的实验表明,在四种模型规模下,一个训练接口适用于固定、混合和动态路由的工作流,并在报告的比较中取得了最强的宏平均结果。优化诊断进一步表征了其在不同事件缩减和集合大小下的稳定性。

英文摘要

Multi-agent systems enable complex reasoning and tool use by coordinating agents that divide roles and refine candidate solutions. Existing methods typically update individual agent responses or treat a complete trajectory as one training example. However, these methods may produce misleading policy updates because they assign the same final outcome to responses or trajectory segments that may play different roles in different team decisions. This is because treating each response as an independent update may separate outputs that jointly determine the next action, while treating an entire trajectory as one update may combine decisions made after different observations. These limitations call for a policy update defined at the level of a team decision, outputs that lead to the same state transition are optimized under a shared objective. In this paper, we propose Setwise Relative Policy Optimization (SRPO) for multi-agent systems. We represent the outputs used together to produce one state transition as an active set. A singleton set covers division of labor, while a larger set covers joint co-evolution, and the set composition can change across decisions. SRPO assigns a shared advantage to each active set, clips the combined policy change, and normalizes its scale according to the set size. This ties each update to the decision that produced the next state. Experiments on mathematical reasoning and multi-turn search demonstrate the effectiveness of SRPO across both tasks. Additional ablation studies analyze the normalization choice and training behavior under changing active-set sizes.

发表机构

  • Southeast University(东南大学)
  • Kuaishou Technology(快手科技)

机构由 AI 辅助整理,请以论文原文为准。

↑