arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向分布式MoE训练的内存高效专家路由

Memory-Efficient Expert Routing for Distributed MoE Training

Arnab Kanti Tarafder, Jaume Guasch-Martí, Gokcen Kestor, Jie Ren

arXiv 2610.07333首次发表:更新:

发表机构

William & Mary; Barcelona Supercomputing Center(威廉与玛丽学院; 巴塞罗那超级计算中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对分布式MoE训练中内存瓶颈,提出基于环的RelayMoE执行模型,通过避免完整top-k扩展缓冲区并重叠通信与计算,实现内存高效路由,在30B-57B模型上取得最高2.02倍吞吐提升和2.85倍序列长度扩展。

AI 中文摘要

随着混合专家(MoE)模型扩展到数百个专家和更高的top-$k$路由,分布式训练中的内存效率成为关键瓶颈。峰值内存主要由MoE块而非注意力机制决定:MoE分发流水线中的每个中间缓冲区都单独按top-k路由缩放。标准的全对全分发器在单个集合通信步骤中发送所有被路由的token,要求一次性构建完整的top-$k$扩展缓冲区。在这项工作中,我们提出了RelayMoE,一种基于环的MoE执行模型,它在专家权重或token循环时进行本地计算,避免了完整的top-$k$扩展分发缓冲区。RelayMoE根据通信量在专家路由和token路由之间进行选择,并将传输与计算重叠。环结构自然支持反向传播期间内存高效的MoE重计算:每一跳重建专家中间结果,使用它们计算梯度,并在下一跳之前释放它们。节省的内存支持更长的序列和更大的批次,或保留更多的注意力激活以减少注意力重计算并提高训练吞吐量。我们在30B至57B的生产级MoE模型和多种专家配置上评估了RelayMoE。在单层MoE实验中,RelayMoE相比Megatron-LM实现了平均2倍加速。在相同GPU内存预算下的全模型训练中,它提高了高达2.02倍的吞吐量,并将最大可测试训练序列长度延长了高达2.85倍。

英文摘要

As Mixture-of-Experts (MoE) models scale toward hundreds of experts and higher top-$k$ routing, memory efficiency in distributed training becomes a critical bottleneck. Peak memory is dominated by the MoE block, not attention: every intermediate buffer in the MoE dispatch pipeline is individually scaled by top-k routing. The standard all-to-all dispatcher sends all routed tokens in a single collective step, requiring the full top-$k$-expanded buffer to be constructed at once. In this work, we propose RelayMoE, a ring-based MoE execution model that computes locally as expert weights or tokens circulate, avoiding full top-$k$-expanded dispatch buffers. RelayMoE selects between expert and token routing according to communication volume and overlaps transfers with computation. The ring structure naturally supports memory-efficient MoE recomputation during backward: each hop reconstructs expert intermediates, uses them to compute gradients, and releases them before the next hop. The saved memory supports longer sequences and larger batches, or retains more attention activations to reduce attention recomputation and improve training throughput. We evaluate RelayMoE on 30B$-$57B production MoE models and varied expert configurations. In single-layer MoE experiments, RelayMoE achieves a $2\times$ average speedup over Megatron-LM. In full-model training under the same GPU memory budget, it improves throughput by up to $2.02\times$ and extends the largest tested trainable sequence length by up to $2.85\times$.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑