结构化MoE专家选择以支持智能体强化学习
Structuring MoE Expert Selection for Agentic Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
针对长视界LLM智能体中的MoE模型,提出分层路由控制框架,通过对齐回合级专家选择与智能体操作并正则化令牌级选择,结合熵门控机制,在多个基准上实现成功率提升超10个百分点。
中文摘要 AI 辅助
长视界LLM智能体常采用稀疏混合专家(MoE)模型实现,然而智能体行为与MoE结构的协同设计仍未得到充分探索。在本工作中,我们全面研究了智能体后训练与MoE专家选择之间的联系。在现成的MoE模型中,我们观察到专家选择展现出一种专门化结构,该结构自然地与智能体轨迹对齐。具体而言,当智能体在不同回合中执行语义相似的操作(如读取、更新)时,专家路由的重叠程度高于执行不同操作时的回合。然而,标准强化学习算法忽略了这种专门化,允许MoE路由在训练过程中不受控制,这在经验上限制了任务性能和推理效率。为解决此问题,我们引入了一种面向智能体任务的分层路由控制框架。我们明确鼓励回合级专家选择与智能体操作对齐,同时正则化令牌级专家选择以保持局部一致性。为解决所提方法在后训练中出现的稳定性问题,我们进一步引入了一种熵门控控制机制。总体而言,我们的路由控制框架在所有评估基准上实现了超过10个百分点的成功率提升。这些结果表明,智能体轨迹结构为在RL后训练期间优化MoE容量提供了有效信号。
英文摘要
Long-horizon LLM agents are frequently implemented using sparse mixture-of-experts (MoE) models, yet the co-design of agentic behavior and MoE structures remains underexplored. In this work, we comprehensively study the connections between agentic post-training and MoE expert selection. In off-the-shelf MoE models, we observe expert selection exhibits a specialized structure that naturally aligns with agentic trajectories. Specifically, expert routing overlaps more between turns where the agent performs semantically similar operations (e.g., READ, UPDATE) than between turns with differing operations. However, standard RL algorithms ignore this specialization, allowing the MoE routing to go uncontrolled during training, which empirically limit task performance and inference efficiency. To address this, we introduce a hierarchical routing control framework for agentic tasks. We explicitly encourage turn-level expert selections to align with agentic operations while regularizing token-level expert selections to maintain local consistency. To resolve stability issues that arise during post-training with the proposed methods, we further introduce an entropy-gated control mechanism. Overall, our routing control framework achieves over 10-point improvements in success rate on all evaluated benchmarks. These results demonstrate that agentic trajectory structure provides an effective signal for optimizing MoE capacity during RL post-training.
发表机构
- Apple(苹果公司)
- Purdue University(普渡大学)
机构由 AI 辅助整理,请以论文原文为准。