arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CIPHER-MoE:在万亿级MoE训练中平衡效率与路由保真度

CIPHER-MoE: Balancing Efficiency and Routing Fidelity in Trillion-Scale MoE Training

Jing Li, Jian Meng, Yingmeng Gao, Suming Qiu, Linyuan Qiu, Dongfang Li, Baotian Hu, Binfan Zheng, Rongqian Zhao, Weijian Sun, Xin Chen

arXiv 2610.05744首次发表:更新:

发表机构

Tongji University; Cornell University; Harbin Institute of Technology, Shenzhen; AI Training Platform Team, Shenzhen Loop Area Institute(同济大学; 康奈尔大学; 哈尔滨工业大学(深圳); 深圳环域研究院AI训练平台团队)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对万亿级MoE训练中的工作负载不平衡问题,提出CIPHER-MoE方法,通过亲和感知的专家到令牌过滤和显式容量控制,在不改变路由选择的前提下减少热点专家负载,实现高达64.9%的负载降低和1.94倍训练加速。

AI 中文摘要

混合专家(MoE)已被广泛应用于近期的大语言模型(LLM)架构中。然而,在LLM训练中扩展MoE引入了系统层面的训练挑战,其中非均匀的令牌路由可能导致专家和设备之间的工作负载高度不平衡,进一步破坏训练过程的稳定性。对于万亿级LLM,不平衡的专家工作负载进一步放大了MoE训练的资源成本,导致负载不足的专家训练效率和硬件利用率下降,而热门专家则需要额外资源来容纳过量的工作负载。近期研究通过复杂的并行策略或资源重新分配来解决不平衡的MoE训练问题。然而,这些系统级方法往往引入额外的资源需求和相当大的编排复杂性,在计算资源受限的情况下训练万亿参数LLM时,这些成本变得越来越难以承受。本工作提出了CIPHER-MoE,它在保持路由器令牌侧Top-K选择不变的同时缓解了MoE工作负载不平衡。CIPHER-MoE应用了具有显式容量控制的亲和感知专家到令牌过滤,以减少热点专家工作负载,而无需额外的硬件资源或复杂的运行时设计。该方法已在包括DeepSeek-V4-Pro在内的大规模MoE模型上进行了评估,实现了高达64.9个百分点的Top-1专家工作负载减少和1.10倍至1.94倍的训练加速,同时保持了训练质量。源代码即将发布。

英文摘要

Mixture-of-Experts (MoE) has been widely adopted in recent large language model (LLM) architectures. However, scaling up MoE in LLM training introduces system-level challenges on training, where non-uniform token routing can lead to highly imbalanced workloads across experts and devices, further destabilizing the training process. With trillion-scale LLMs, imbalanced expert workloads further amplify the resource cost of MoE training, resulting in degraded training efficiency and hardware utilization for underloaded experts, while hot experts require additional resources to accommodate excessive workloads. Recent studies address imbalanced MoE training through intricate parallelism strategies or resource reallocation. However, these system-level approaches often introduce additional resource requirements and considerable orchestration complexity, which become increasingly difficult to afford when training trillion-parameter LLMs under constrained computational resources. This work introduces CIPHER-MoE, which mitigates MoE workload imbalance while keeping the router's token-side Top-K selection unchanged. CIPHER-MoE applies affinity-aware Expert-to-Token filtering with explicit capacity control to reduce hotspot expert workloads without additional hardware resources or complex runtime design. The proposed method has been evaluated on large-scale MoE models, including DeepSeek-V4-Pro, showing up to 64.9 percentage points Top-1 expert workload reduction and 1.10$\times$-1.94$\times$ training acceleration, while preserving the training quality. The source code will be released soon.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑