发表机构
NVIDIA(英伟达)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出CE-MoE模型,通过异构层模式解耦令牌与通道混合深度,在2B至31.5B参数规模下降低训练成本,31.5B时减少33.3%GPU小时数,下游性能与推理效率均提升。
AI 中文摘要
在使用专家并行性训练混合专家(MoE)语言模型时,全对全(all-to-all)的令牌分发与收集通信操作会消耗大量端到端训练时间。本研究探讨通信高效MoE模型(CE-MoE),采用异构层模式,将令牌混合与通道混合的深度解耦。与传统模型(每一层令牌混合层后插入MoE层,如注意力机制、Mamba-2)相比,CE-MoE模型将专家容量集中在少数选定的路由MoE层中,同时通过添加额外的令牌混合层和密集前馈网络(dense-FFN)层维持模型深度。在总参数规模从20亿到315亿的缩放阶梯上,在总参数与激活参数匹配的条件下,CE-MoE模型在训练成本持续降低的同时,验证损失与全MoE基线相当,下游基准测试性能也匹配。在315亿参数规模下,CE-MoE减少了33.3%的GPU小时数,且平均下游得分与推理吞吐量均有所提升。
英文摘要
When training Mixture-of-Experts (MoE) language models with expert parallelism, all-to-all token dispatch and combine collectives can consume a substantial fraction of end-to-end training time. In this work, we study communication-efficient MoE models (CE-MoE), in which we adopt a heterogeneous layer pattern that decouples token-mixing and channel-mixing depth. Compared to conventional models which interleave MoE layers after each token-mixing layer (e.g., attention, Mamba-2), CE-MoE models concentrate expert capacity in a select few routed MoE layers, while maintaining depth by adding additional token-mixing and dense-FFN layers. Across a scaling ladder from 2B to 31.5B total parameters, under matched total and activated parameters, CE-MoE models consistently reduce training cost while matching validation loss and downstream benchmarks with full-MoE baselines. At the 31.5B scale, CE-MoE uses 33.3\% fewer GPU-hours while improving average downstream score and inference throughput.