arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39137cs.LG

ID Balancing:基于PID负载控制的极稀疏MoE稳定训练

ID Balancing: Stable Training of Extremely Sparse MoE via PID-Based Load Control

Peng Jin, Zihan Qiu, Zekun Wang, Bo Zheng, Yang Xu, Tian Xie, Xiao Li, Huaqing Zhang, Haoran Lian, Rui Men, Dayiheng Liu

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出ID Balancing,一种基于PID控制的积分-微分负载均衡方法,用于稳定训练极稀疏MoE模型,显著降低负载不均衡并提升扩展性能。

中文摘要 AI 辅助

通过混合专家(MoE)扩展大型语言模型(LLM)能够在每token计算量几乎不变的情况下实现参数规模的巨大增长。然而,进一步扩大参数数量需要越来越稀疏的路由,此时专家负载不均衡问题变得更加严重。这种不均衡会降低参数利用率和训练效率,并可能损害训练稳定性,成为可靠扩展的瓶颈。在本工作中,我们将两种具有代表性的免辅助损失方法统一为不完整的比例-积分-微分(PID)控制器:DeepSeek的无损失方法充当固定步长的积分控制器,而Kimi K3的分位数平衡则充当广义比例控制器。基于这一控制视角,我们提出了ID Balancing,一种积分-微分控制器。它根据负载误差缩放积分项,并仅在负载不均衡恶化时激活微分项,从而对较大或恶化的误差进行更强修正,并在接近平衡时进行较小更新。在768个专家上的Top-10、Top-5和Top-3路由评估中,与Top-3设置下的最佳基线相比,ID Balancing将最坏情况下的骨干网络MaxVio和训练平均骨干网络MinVio分别降低了超过50%和12%。当总参数数量从18.9B增加到69.9B(768个专家中的Top-10)时,ID Balancing的最坏情况骨干网络MaxVio几乎保持不变,并且比辅助损失基线低约89.6%。ID Balancing还保持了具有竞争力的语言建模和下游任务性能。随着稀疏度的增加,ID Balancing的优势愈发明显,使其成为扩展更大、更稀疏MoE模型的有前景的解决方案。

英文摘要

Scaling Large Language Models (LLMs) via Mixture-of-Experts (MoE) enables massive parameter growth with nearly constant per-token computation. However, further scaling the parameter count requires increasingly sparse routing, where expert load imbalance becomes more severe. This imbalance reduces parameter utilization and training efficiency, and can undermine training stability, becoming a bottleneck to reliable scaling. In this work, we unify two representative auxiliary-loss-free methods as incomplete Proportional-Integral-Derivative (PID) controllers: DeepSeek's loss-free method acts as a fixed-step integral controller, while Kimi K3's Quantile Balancing functions as a generalized proportional controller. Building on this control perspective, we propose ID Balancing, an Integral-Derivative controller. It scales its integral term with load error and activates its derivative term only when imbalance worsens, enabling stronger corrections for large or worsening errors and smaller updates near balance. Evaluated across Top-$10$, Top-$5$, and Top-$3$ routing over $768$ experts, ID Balancing reduces worst-case backbone MaxVio and training-average backbone MinVio by over $50\%$ and $12\%$, respectively, relative to the best baselines in the Top-$3$ setting. When the total parameter count increases from $18.9$B to $69.9$B (Top-$10$-of-$768$), ID Balancing's worst-case backbone MaxVio remains nearly unchanged and is approximately $89.6\%$ lower than that of the auxiliary-loss baseline. ID Balancing also maintains competitive language-modeling and downstream performance. The advantages of ID Balancing grow as sparsity increases, making it a promising solution for scaling larger, sparser MoE models.

发表机构

  • Qwen Team, Alibaba Token Hub, Alibaba Group(通义千问团队,阿里巴巴令牌中心,阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

↑