arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一个模型,多种心智:通过角色混合在单个智能体中解锁多智能体协同

One Model, Many Minds: Unlocking Multi-Agent Synergy in a Single Agent via Mixture of Roles

Zhichen Zeng, Huiyuan Chen, Jingru Cheng, Juan Zha, Ming Liu, Ying Chen, Xiyuan Yang, Chaosheng Dong, Haiyang Zhang, Hanghang Tong

arXiv 2608.27338首次发表:更新:

AI 中文总结

该研究提出 MoRe 方法,将多种专门化组合为单个引导向量,使单智能体实现多视角专门化,性能优于单智能体基线,与 MAS 相当且 token 成本大幅降低。

AI 中文摘要

将大型语言模型(LLM)专门化为不同能力,支撑了从个性化助手到多智能体系统(MAS)的各类成功应用。单智能体范式依赖预定义的 persona(角色设定)或引导向量来实现专门化,但它们施加单一固定的专门化,无法适配多样化查询。相反,MAS 通过协调具有不同文本角色的智能体实现动态多视角问题解决,但融合这些专门化需要多轮交互,这会增加上下文长度和推理成本。为解决这些局限,我们提出角色混合(MoRe),它将多种专门化自适应组合为单个引导向量,用于单轮推理。具体而言,MoRe 学习多样化的引导向量码本,每个向量编码一个潜在角色;查询感知路由器动态将码本融合为包含多个角色的引导向量。通过用组合向量引导主干 LLM,MoRe 在单智能体、单轮推理过程中实现多视角专门化。所提出的 MoRe 可通过三阶段 SFT(监督微调)课程和 GRPO(生成式偏好优化)后训练高效训练,且主干 LLM 保持冻结。在推理和人格基准上的实验表明,MoRe 比单智能体基线平均性能提升 2.2%,达到与 MAS 相当的性能,同时 token 成本降低 20 倍。

英文摘要

Specializing Large Language Models (LLMs) toward distinct abilities underpins successes ranging from personalized assistants to multi-agent systems (MAS). Single-agent paradigms rely on pre-defined personas or steering vectors to induce specialization, yet they impose a single fixed specialization that fails to adapt to diverse queries. Conversely, MAS achieves dynamic multi-perspective problem solving by orchestrating agents with distinct text-based roles, but fusing these specializations requires multi-turn interactions that inflate context length and inference cost. To address these limitations, we propose Mixture of Roles (MoRe), which adaptively composes multiple specializations into a single steering vector for single-turn inference. Specifically, MoRe learns a diversified codeboox of steering vectors, each of which encodes a latent role. A query-aware router dynamically fuses the codebook into a steering vector that encompasses multiple roles. By steering the backbone LLM with the composed vector, MoRe enables multi-perspective specialization in a single-agent, single-turn inference process. The proposed MoRe can be efficiently trained via a three-stage SFT curriculum and GRPO post-training, while the backbone LLM remains frozen. Experiments across reasoning and personality benchmarks show that MoRe outperforms single-agent baselines by 2.2% on average, and achieves performance on par with MAS while reducing token cost by 20x.

Comments19 pages, 10 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑