arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MoE预训练中的专家耦合:利用关联放置与令牌洗牌降低全对全通信开销

Expert Coupling in MoE Pretraining: Reducing All-to-All Overhead with Correlated Placement and Token Shuffling

Radha Gulhane, Quentin Anthony, Beren Millidge

arXiv 2610.09372首次发表:更新:

发表机构

Zyphra(Zyphra)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出关联专家放置与令牌洗牌两种方法,利用预训练中路由器学到的令牌-专家关联模式,减少MoE专家并行下的全对全通信开销,最高可将端到端步骤时间减少1.41倍,且不改变路由决策与专家参数。

AI 中文摘要

混合专家(MoE)层将Transformer的前馈模块替换为E个专家网络,每个令牌被路由到其中的k个专家。在专家并行(EP)下,专家分布在多个GPU上,每个MoE层在前向和后向传播中执行全对全集合通信,以将令牌分发给其专家并合并结果。在每节点配备8个AMD Instinct MI300X GPU的集群上,当EP度为32且采用top-2路由时,这些集合通信占训练步骤时间的45%;采用top-6路由时,占比达到60%。我们发现,在预训练早期,路由器已经学会以关联模式将令牌分配给专家,这种关联既存在于同一层内,也存在于跨层之间。在top-2路由下,一个层中0.8%的专家对被42%的令牌共同选择,且令牌在某层选择的专家能预测其在下一层选择的专家。我们利用这些关联性将更多的令牌-专家分配保留在令牌所在的GPU上,从而减少跨GPU和跨节点的通信。关联专家放置将经常被一起选择的专家放在同一GPU上。结合一个将每个令牌仅发送到每个GPU一次的调度器,该方法最多可减少58%的调度行数。令牌洗牌适用于序列并行将令牌分片到EP组的情况。它在注意力之后的归约散播过程中,将每个令牌移动到预测持有其下一层专家的GPU上。在单节点上,这将令牌-专家分配在令牌所在GPU上处理的比例从12.5%提升至59%。在Megatron-LM中,当EP度从8到64且采用top-2和top-6路由时,这两种方法将全对全通信时间减少了1.16-2.63倍,端到端步骤时间最多减少1.41倍。这两种方法均不改变模型底层的路由决策或专家参数。

英文摘要

Mixture-of-Experts (MoE) layers replace the feed-forward block of a Transformer with E expert networks, and each token is routed to k of these experts. Under expert parallelism (EP) the experts are distributed across GPUs, and every MoE layer runs all-to-all collectives in the forward and backward passes to dispatch tokens to their experts and then combine the results. On a cluster with 8 AMD Instinct MI300X GPUs per node, these collectives can take 45% of the training step at EP32 with top-2 routing and 60% with top-6 routing. We find that early in pretraining routers have already learned to assign tokens to experts in correlated patterns, both within a layer and across layers. At top-2, 0.8% of the expert pairs in a layer are selected together by 42% of tokens, and the experts a token selects at one layer predict the experts it selects at the next layer. We use these correlations to keep more token--expert assignments on the token's own GPU, which reduces communication across GPUs and across nodes. Correlated expert placement puts experts that are often selected together on the same GPU. Combined with a dispatcher that sends each token to each GPU once, it removes up to 58% of dispatched rows. Token shuffling applies when sequence parallelism shards tokens across the EP group. It moves each token to the GPU predicted to hold its next-layer experts during the reduce-scatter that follows attention. On one node this raises the share of token--expert assignments served on the token's GPU from 12.5% to 59%. In Megatron-LM, across EP degrees from 8 to 64 with top-2 and top-6 routing, the two methods reduce all-to-all time by 1.16-2.63X and end-to-end step time by up to 1.41X. Neither method changes the models' underlying routing decisions or expert parameters.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑