Cobalt:利用专家共激活实现高效的分布式MoE训练
Cobalt: Leveraging Expert Co-activation for Efficient Distributed MoE Training
浏览论文内容
中文总结 AI 辅助
Cobalt提出利用专家共激活优化MoE训练,通过两阶段布局规划和通信感知路由,在32个B200 GPU上实现最高2.41倍加速并减少99.26%跨节点流量。
中文摘要 AI 辅助
混合专家(MoE)模型已成为扩展大型语言模型的主流方法,因为它能在保持计算成本几乎不变的同时扩大模型容量。训练大规模MoE模型依赖于专家并行(EP),该策略将专家副本分布到各GPU上,并通过全对全通信交换令牌。EP的效率常受两个系统瓶颈制约:跨节点令牌传输受限于节点间带宽,而不均衡的专家工作负载导致GPU间计算不平衡。先前的工作基于每个专家的工作负载统计来缓解这些瓶颈,但忽略了专家之间可能共享通信这一事实。在本工作中,我们通过实验观察到,许多专家对经常被单个令牌共同激活。基于此,我们提出了Cobalt,一个高效的MoE训练框架,利用专家共激活来减少跨节点流量和工作负载不均衡。Cobalt采用两阶段专家布局规划器,根据不断变化的专家共激活和工作负载条件调整专家布局。它定期将频繁共激活的专家共置于同一节点以减少跨节点通信,并执行每步的节点内调整以重新平衡工作负载。随后,我们开发了一种通信感知的任务分配方法,根据当前专家布局将令牌路由到更少的远程节点。在32个B200 GPU上的实验表明,与现有MoE训练框架相比,Cobalt实现了最高1.53-2.41倍(平均1.28-1.89倍)的加速,同时将跨节点令牌流量减少了75.74%-99.26%。
英文摘要
Mixture-of-Experts (MoE) has increasingly become a mainstream approach for scaling large language models, as it expands model capacity while keeping computation cost nearly constant. Training large-scale MoE models relies on Expert Parallelism (EP), which distributes expert replicas across GPUs and exchanges tokens through all-to-all communication. The efficiency of EP is often constrained by two system bottlenecks: cross-node token transfers are limited by inter-node bandwidth, while skewed expert workloads lead to imbalanced computation across GPUs. Prior work mitigates these bottlenecks based on per-expert workload statistics, but overlooks the fact that experts could share the communication. In this work, we empirically present the observation that many pairs of experts are frequently co-activated by individual tokens. Motivated by this, we present Cobalt, an efficient MoE training framework that leverages expert co-activation to reduce cross-node traffic and workload imbalance. Cobalt adopts a two-stage expert layout planner that adapts expert layout to the evolving expert co-activation and workload conditions. It periodically co-locates frequently co-activated experts on the same node to reduce the cross-node communication, and performs per-step intra-node adjustment to rebalance the workloads. Subsequently, we develop a communication-aware task assignment method that routes tokens to fewer remote nodes based on the current expert layout. Experiments on 32 B200 GPUs show that Cobalt achieves up to 1.53-2.41 times (1.28-1.89 times on average) of speedup compared to existing MoE training frameworks, while reducing cross-node token traffic by 75.74%-99.26%.
发表机构
- Zhejiang University(浙江大学)
- Peking University(北京大学)
- Shanghai Jiao Tong University(上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。