发表机构
University of Science and Technology of China(中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对专家并行MoE训练中的负载不均衡问题,提出GPU原生、拓扑感知的负载均衡系统TopoEP,通过确定性GPU求解器生成热门专家复制与令牌重路由计划,在32-GPU集群上提升训练吞吐量6.2%至11.4%。
AI 中文摘要
动态路由在大规模专家并行混合专家(MoE)训练中造成严重的负载不均衡,使得承载热门专家的GPU成为掉队者。由于每个MoE层都要等待其最慢的秩,这些掉队者延长了专家并行阶段并降低了整体训练效率。现有的专家并行负载均衡(EPLB)系统通常在CPU上计算负载均衡计划,导致设备-主机数据传输和跨秩同步,使得在每一层和每个微批次上进行调度代价高昂。其规划方案还忽视了现代扩展(scale-up)和扩展(scale-out)GPU集群的层级通信成本。我们提出了TopoEP,一个面向大规模MoE训练的GPU原生、拓扑感知的负载均衡系统。在每个MoE层和训练微批次上,TopoEP将当前路由结果转换为热门专家复制和令牌重路由决策,并在无需数据依赖的主机同步的情况下执行所得计划,从而减少关键路径开销。为了生成这些决策,TopoEP使用一个确定性GPU求解器,先执行节点间放置,再进行节点内细化,使所有秩能够独立生成逐位相同的计划。在32-GPU NVIDIA H800集群上,将TopoEP与Megatron-LM集成,在三个代表性MoE模型上将端到端训练吞吐量提升了6.2%至11.4%。
英文摘要
Dynamic routing creates severe load imbalance in large-scale expert-parallel Mixture-of-Experts (MoE) training, turning GPUs that host hot experts into stragglers. As each MoE layer waits for its slowest rank, these stragglers prolong the expert-parallel stage and reduce overall training efficiency. Existing expert-parallelism load-balancing (EPLB) systems commonly compute load-balancing plans on the CPU, incurring device--host data transfers and cross-rank synchronization that make scheduling at every layer and microbatch expensive. Their planning formulations also overlook the hierarchical communication costs of modern scale-up and scale-out GPU clusters. We present \textit{TopoEP}, a GPU-native, topology-aware load-balancing system for large-scale MoE training. At each MoE layer and training microbatch, \textit{TopoEP} converts the current routing result into hot-expert replication and token-rerouting decisions and executes the resulting plan without data-dependent host synchronization, reducing critical-path overhead. To generate these decisions, \textit{TopoEP} uses a deterministic GPU solver that performs inter-node placement followed by intra-node refinement, allowing all ranks to independently produce bitwise-identical plans. On a 32-GPU NVIDIA H800 cluster, integrating \textit{TopoEP} with Megatron-LM improves end-to-end training throughput by 6.2\%--11.4\% across three representative MoE models.