发表机构
Princeton University; NVIDIA(普林斯顿大学; 英伟达)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MegaFlux通过运行时专家复制和流水线通信,在MoE超级内核中缓解路由偏斜,实现前向1.45倍、后向1.28倍的几何平均加速。
AI 中文摘要
混合专家(MoE)超级内核将专家并行通信与专家计算融合在一起。然而,在固定专家放置下,路由偏斜会造成GPU掉队者:过载的GPU决定层延迟,而其他GPU则空闲。复制热门专家可以将工作转移到负载不足的GPU上,但动态复制会引入额外工作:复制品必须接收专家权重才能执行,并且在训练期间,其部分权重梯度必须在专家所有者处进行归约。我们提出MegaFlux,它将专家复制变为运行时决策,并在持久化MoE执行中流水线化复制所引发的通信。一个设备端规划器在每GPU复制预算下联合选择复制位置并分配瓦片对齐的令牌块,保持路由器输出不变。前向和后向超级内核实现了流水线专家复制:复制品在所需权重到达时开始计算,而后向将复制品梯度归约与正在进行的专家计算重叠。MegaFlux扩展了TensorRT-LLM的CuTeDSL MegaMoE前向内核,并引入了新的后向MoE超级内核。在八块NVIDIA B200 GPU上,每个方向147种配置中,MegaFlux相对于固定放置的相同超级内核,前向几何平均加速比为1.45倍,后向为1.28倍,峰值分别达到2.14倍和2.64倍。在消融实验中,流水线隐藏了前向中56%–76%的复制品权重传输成本,以及后向中91%–100%的权重传输和复制品梯度归约组合成本,相比相同复制计划但分开执行这些操作,额外实现了高达13.2%和26.7%的层延迟降低。集成到vLLM用于DeepSeek-V4-Pro预填充时,MegaFlux相对于固定放置提供了1.13–1.26倍的中位端到端加速比。
英文摘要
Mixture-of-experts (MoE) megakernels fuse expert-parallel communication with expert computation. However, under fixed expert placement, routing skew creates GPU stragglers: overloaded GPUs determine layer latency while others sit idle. Replicating hot experts can shift work to underloaded GPUs, but dynamic replicas introduce additional work: replicas must receive expert weights to execute and, during training, their partial weight gradients must be reduced at the expert owners. We present MegaFlux, which makes expert replication a runtime decision and pipelines the communication induced by replication within persistent MoE execution. An on-device planner jointly selects replica locations and assigns tile-aligned token blocks under a per-GPU replica budget, leaving router outputs unchanged. The forward and backward megakernels realize pipelined expert replication: replicas begin computation as their required weights arrive, while backward overlaps replica-gradient reduction with ongoing expert computation. MegaFlux extends TensorRT-LLM's CuTeDSL MegaMoE forward kernel and introduces a new backward MoE megakernel. Across 147 configurations per direction on eight NVIDIA B200 GPUs, MegaFlux achieves geometric-mean speedups of $1.45\times$ for forward and $1.28\times$ for backward over the same megakernels with fixed placement, peaking at $2.14\times$ and $2.64\times$. In ablations, pipelining hides $56$--$76$% of replica-weight transfer cost in forward and $91$--$100$% of combined weight-transfer and replica-gradient-reduction cost in backward, yielding up to $13.2$% and $26.7$% additional layer-latency reductions over the same replication plans with these operations executed separately. Integrated into vLLM for DeepSeek-V4-Pro prefill, MegaFlux delivers $1.13$--$1.26\times$ median end-to-end speedups over fixed placement.
Comments15 pages, 7 figures, 1 table