发表机构
Shanghai AI Laboratory; Shanghai Jiao Tong University; Nanyang Technological University(上海人工智能实验室; 上海交通大学; 南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对稀疏混合专家模型静态Top-k专家选择导致的脆弱边界问题,提出弹性专家路由方法,在两类训练场景下均提升了模型下游任务性能。
AI 中文摘要
稀疏混合专家(Sparse Mixture-of-Experts)模型在保持每个token固定计算预算的同时,能高效扩展参数容量。然而,传统训练范式强制采用静态的Top-k专家选择,将连续的路由分布转化为刚性阶跃函数,这种约束会产生脆弱的边界:基于微小的分数波动,竞争力极强的专家被随意划分到全监督区域和零反馈区域。为解决该问题,我们提出弹性专家路由(Elastic Expert Routing),该方法从以k为中心的局部离散分布中随机采样活跃专家预算。在多个训练迭代中,此机制将尖锐阈值柔化为渐进概率分布。由于采样邻域保持对称,该方法与确定性训练的预期计算成本匹配,同时保留推理预算。大量实验在监督微调与从头预训练两种场景下验证了该方法的有效性:在监督微调中,弹性路由使OLMoE-1B-7B和Qwen3-30B-A3B的下游宏平均分别提升0.84和2.02个百分点;在从头预训练中,其在下游任务上的平均表现比静态Top-k基线高出1.6个百分点。
英文摘要
Sparse Mixture-of-Experts models scale parameter capacity efficiently while maintaining a fixed compute budget per token. However, traditional training paradigms enforce a static choice of top-$k$ experts, which converts a continuous routing distribution into a rigid step function. This constraint introduces a brittle boundary where highly competitive experts are arbitrarily separated into full-supervision and zero-feedback zones based on minor score fluctuations. To address this issue, we propose Elastic Expert Routing, which stochastically samples the active expert budget from a localized discrete distribution centered at $k$. Over multiple training iterations, this mechanism softens the sharp threshold into a gradual probability distribution. Because the sampling neighborhood remains symmetric, this approach matches the expected computational cost of deterministic training, while preserving the inference budget. Extensive experiments demonstrate the efficacy of our method on both supervised fine-tuning and from-scratch pretraining settings. During supervised fine-tuning, elastic routing improves downstream macro-averages on OLMoE-1B-7B and Qwen3-30B-A3B by $+0.84$ and $+2.02$ points, respectively. In addition, in from-scratch pretraining, it outperforms the static top-$k$ baseline by $1.6$ points on average across downstream tasks.
Comments13 pages, 4 figures