arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.28053cs.LGcs.CL

混合专家模型的精确分位数均衡与负载误差注入

Exact Quantile Balancing and Load-Error Injection for Mixture-of-Experts

Pit Neitemeier, Jiaze Li, Alessio Serra, Philipp Scholl, Sohir Maskey

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出精确分位数均衡(EQB)与负载误差注入(LEI),分别改进MoE训练的全局与局部负载均衡,在7.5B参数模型上提升性能并优于GShard损失。

中文摘要 AI 辅助

混合专家(MoE)训练需要全局负载均衡以防止专家利用不足,以及局部均衡以实现高效的专家并行执行。现有的分布式分位数均衡(QB)使用依赖分片或近似的全局分位数,而令牌无关的专家偏置无法确保微批次级别的均衡。我们引入了精确分位数均衡(EQB),它以可忽略的通信开销计算精确的全局批次BF16分位数,以及负载误差注入(LEI),它将局部负载误差直接注入到路由器分数的梯度中。在训练高达5000亿令牌的7.5B参数MoE上,EQB相比朴素QB改善了全局均衡和下游性能,而LEI改善了局部均衡,并在相当质量下优于GShard损失。

英文摘要

Mixture-of-Experts (MoE) training requires global load balance to prevent expert under-utilization and local balance for efficient expert-parallel execution. Existing distributed Quantile Balancing (QB) uses shard-dependent or approximate global quantiles, while token-independent expert biases cannot ensure microbatch-level balance. We introduce Exact Quantile Balancing (EQB), which computes exact global-batch BF16 quantiles with negligible communication, and Load-Error Injection (LEI), which injects local load errors directly into router-score gradients. On 7.5B-parameter MoEs trained for up to 500B tokens, EQB improves global balance and downstream performance over naive QB, while LEI improves local balance and outperforms the GShard loss at comparable quality.

发表机构

  • Aleph Alpha

机构由 AI 辅助整理,请以论文原文为准。

↑