arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32398cs.AI

形式服从功能:基于副本专家机制的混合专家模型中的分布正交化

Function Over Form: Distributional Orthogonalization in Mixture-of-Experts with Replica Expert Mechanism

Jinfan He, Yunzhuo Liu, Kai Zhang, Weidong Han, Key, Rayying

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出分布正交化损失和副本专家机制,通过惩罚专家负载签名重叠防止坍缩并改善负载均衡,在两种MoE模型上优于现有路由算法。

中文摘要 AI 辅助

大语言模型的扩展日益依赖于混合专家(MoE)架构,以将活跃计算与总参数量解耦。然而,MoE的有效性常受限于专家坍缩和表示冗余,这两者均导致模型容量利用不足。为解决这些挑战,本文提出分布正交化损失(DO-loss),一种辅助正则化方法,将关注点从静态权重多样性转向动态路由行为。通过将每个专家的令牌分配历史表示为高维二进制负载签名,DO-loss惩罚签名重叠以防止专家坍缩,同时鼓励功能特化。为使该算法设计与系统效率对齐,我们进一步引入副本专家机制(REM),通过双层策略改善负载均衡:在全局批次级别调整副本专家放置,并在微批次级别执行实时令牌调度。实证评估表明,我们的方法在4.8B A0.5B和30B A3B两种MoE模型的下游任务上均优于所评估的路由算法,同时保持可比的训练效率。

英文摘要

The scaling of LLMs increasingly relies on MoE architectures to decouple active computation from total parameter count. However, the efficacy of MoE is often constrained by expert collapse and representation redundancy, both leading to underutilization of model capacity. To address these challenges, this paper proposes Distributional Orthogonalization Loss (DO-loss), an auxiliary regularization that shifts the focus from static weight diversity to dynamic routing behavior. By representing each expert's token assignment history as a high-dimensional binary load signature, DO-loss penalizes signature overlap to prevent expert collapse while encouraging functional specialization. To align this algorithmic design with system efficiency, we further introduce the Replica Expert Mechanism (REM), which improves load balancing through a two-tiered strategy: adjusting replica expert placement at the global-batch level and performing real-time token dispatching at the micro-batch level. Empirical evaluations demonstrate that our method outperforms the evaluated routing algorithms on downstream tasks for both 4.8BA0.5B and 30BA3B MoE models, while maintaining comparable training efficiency.

补充信息

↑