RoutePack:面向MoE强化学习的专家放置与注意力感知数据打包
RoutePack: Expert Placement and Attention-Aware Data Packing for MoE Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
RoutePack是一种分层规划器,通过协调专家重路由与注意力感知数据打包,在Ling-3.0-Tiny和Ling-3.0-Flash上分别提升MoE强化学习模型吞吐量8.85%和14.89%。
中文摘要 AI 辅助
训练用于强化学习(RL)的混合专家(MoE)模型会耦合两个负载均衡问题:序列组成决定每个数据并行微批次中的密集注意力工作,而令牌路由决定专家并行秩上的稀疏专家工作;仅优化其中一个问题会将瓶颈转移到另一个问题。在MoE RL中,推理时的路由重放在训练步骤前就暴露了每个样本的序列长度和逐层专家需求。我们提出RoutePack,这是一种分层规划器,它在优化器步骤窗口内协调状态一致的逐层专家重路由与注意力及专家感知的联合数据打包。RoutePack首先使用聚合路由需求在每个MoE层独立放置专家;然后将样本打包为最小的经认证或已知可行的令牌上限执行行数量,并通过考虑EDP分片的目标优化其DP布局,该目标结合了窗口归一化的线性二次注意力代理与每层物理EP秩峰值,同时最小化最慢EDP分片的累积成本。并行种群退火在固定行可行布局中搜索,同时保持样本覆盖、容量、非空单元、相等微批次数量和通信器拓扑。状态一致的物化保留了逻辑top-k路由和现有MoE内核,无需微批次级别的专家复制。在Ling-3.0-Tiny和Ling-3.0-Flash上,专家重路由使平均训练器测量的令牌吞吐量分别提升3.80%和10.50%,而路由感知打包又分别增加4.86%和3.98%;总体而言,RoutePack相比基线分别提升吞吐量8.85%和14.89%。
英文摘要
Training Mixture-of-Experts (MoE) models for reinforcement learning (RL) couples two load-balancing problems: sequence composition determines dense attention work in each data-parallel microbatch, while token routing determines sparse expert work on expert-parallel ranks. Optimizing either alone can shift the bottleneck to the other. In MoE RL, rollout-time routing replay exposes every sample's sequence length and layer-wise expert demand before its training step. We present RoutePack, a hierarchical planner that coordinates state-consistent, layer-wise expert rerouting with joint attention- and expert-aware data packing over an optimizer-step window. RoutePack first places experts independently at each MoE layer using aggregate routing demand. It then packs samples into the smallest certified, or best-known feasible, number of token-capped execution rows and optimizes their DP layout with a projected EDP-shard-aware objective. The objective combines a window-normalized linear-quadratic attention proxy with per-layer physical EP-rank peaks and minimizes the accumulated cost of the slowest EDP shard. Parallel population annealing searches fixed-row feasible layouts while preserving sample coverage, capacity, nonempty cells, equal microbatch counts, and communicator topology. State-consistent materialization preserves logical top-k routing and existing MoE kernels without microbatch-level expert replication. Across Ling-3.0-Tiny and Ling-3.0-Flash, expert rerouting improves mean trainer-measured token throughput by 3.80% and 10.50%, while routing-aware packing adds another 4.86% and 3.98%, respectively. Overall, RoutePack improves throughput by 8.85% and 14.89% over the baseline.
发表机构
- Ant Group(蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。