AI 中文总结
针对MoE训练中动态路由引发的负载不平衡问题,提出TAOT拓扑感知最优传输方法,可实现1.43倍训练加速,平衡质量优异且通信成本降幅最高达74%。
AI 中文摘要
混合专家(Mixture-of-Experts,MoE)已成为缩放大型语言模型(Large Language Models,LLMs)的关键架构,但其动态路由会在专家并行训练中引发严重的负载不平衡问题。现有的动态副本方法会将热门专家复制到空闲秩以分担计算负载,但这类方法仅优化负载平衡,忽略了跨多节点拓扑迁移专家权重的成本,导致产生的跨节点通信开销可能超过负载平衡带来的收益,进而推高训练成本。本文提出TAOT,一种面向动态专家副本放置的拓扑感知最优传输方法。TAOT将热门秩的过载情况与低负载秩的空闲容量建模为带通信代价矩阵的平衡熵正则化最优传输问题,通过Sinkhorn-Knopp迭代求解得到秩级流提示,再将整数副本匹配与令牌分配结合为可执行调度方案。在系统层面,TAOT将 guest 权重传输与主专家计算重叠执行,以隐藏通信开销。实验表明,TAOT可实现1.43倍的MoE训练端到端加速,达到与现有最先进方法相当或更优的平衡质量,且在所有配置下均实现最低的加权专家通信成本,降幅最高达74%。
英文摘要
Mixture-of-Experts (MoE) has become a key architecture for scaling large language models (LLMs), yet its dynamic routing causes severe load imbalance in expert-parallel training. Existing dynamic-replica methods copy hot experts onto idle ranks to share computation, but they optimize load balance alone and ignore the cost of moving expert weights across a multi-node topology, so the resulting cross-node communication can outweigh the balancing gain and inflate training cost. We present TAOT, a topology-aware optimal transport method for dynamic expert-replica placement. TAOT models the overload on hot ranks and the spare capacity on lightly loaded ranks as a balanced entropy-regularized optimal transport problem with a communication-cost matrix, solves it with Sinkhorn-Knopp iterations to produce rank-level flow hints, and combines integer replica matching with token assignment into an executable schedule. At the system level, it overlaps guest-weight transfer with home-expert computation to hide the communication overhead. Experiments show TAOT achieves a 1.43x end-to-end MoE training speedup, reaches balance quality competitive with or better than existing state-of-the-art methods, and attains the lowest weighted expert-communication cost across all configurations, with up to a 74% reduction.