发表机构
New York University; Center for Data Science, NYU Shanghai(纽约大学; 上海纽约大学数据科学中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对MoE路由不完善导致专家训练不足的可靠性问题,提出分布鲁棒MoE训练(DRMoET),通过熵正则化softmax优化高损失路由,在10.3B规模下将七任务平均从0.6625提升至0.6767,并降低专家损失方差。
AI 中文摘要
专家混合(MoE)Transformer通过每个token仅激活少数专家来扩展容量,但这种稀疏性带来了一个隐藏的可靠性问题:当路由不完美时,负载均衡模型可能将token发送到对分配输入训练不足的专家。我们提出分布鲁棒MoE训练(DRMoET),这是一种即插即用的目标函数,将逐层专家视为内生鲁棒性组,并优化高损失路由结果,而非仅仅均衡流量。DRMoET通过EMA平滑、激活加权的专家损失上的熵正则化softmax规则更新每层专家分布,强化可行的非最优路由路径,同时保持标准MoE计算。在FLAME-MoE配方下,总参数746M和10.3B规模,DRMoET在下游平均性能上优于标准FLAME-MoE和无辅助损失均衡。在总参数10.3B和67B训练token下,DRMoET将七任务平均从0.6625提升至0.6767,而无辅助损失基线达到0.6431。机制分析显示专家损失方差更低,平均损失几乎不变,在强制中间k错路由下超额损失降低4.3%,并改善了领域专家专业化。这些结果将路由鲁棒性——不仅仅是利用率均衡——定位为可靠稀疏MoE扩展的实际目标。项目页面和代码可在以下网址获取:this https URL。
英文摘要
Mixture-of-Experts (MoE) transformers scale capacity by activating only a few experts per token, but this sparsity creates a hidden reliability problem: when routing is imperfect, load-balanced models may send tokens to experts that are insufficiently trained for the assigned inputs. We propose Distributionally Robust MoE Training (DRMoET), a drop-in objective that treats layer-wise experts as endogenous robustness groups and optimizes high-loss routing outcomes rather than merely equalizing traffic. DRMoET updates a per-layer expert distribution by an entropy-regularized softmax rule on EMA-smoothed, activation-weighted expert losses, strengthening plausible non-top routing paths while preserving standard MoE computation. Under the FLAME-MoE recipe at 746M-total and 10.3B-total scales, DRMoET improves downstream averages over both standard FLAME-MoE and auxiliary-loss-free balancing. At 10.3B total parameters and 67B training tokens, DRMoET improves the seven-task average from 0.6625 to 0.6767, while the auxiliary-loss-free baseline achieves 0.6431. Mechanistic analyses show lower expert-loss variance with nearly unchanged mean loss, 4.3% lower excess loss under forced mid-$k$ misrouting, and improved domain-expert specialization. These results position routing robustness-not only utilization balance-as a practical objective for reliable sparse MoE scaling. Project page and code are available at: https://drmoet.github.io/.
CommentsIn proceedings of NeurIPS 2026