arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32581cs.CLcs.LG

HERO-MoE:具有尺度保持融合的历史专家路由

HERO-MoE: Historical Expert Routing with Scale-Preserving Fusion

Junxiang Qiu, Zhengsu Chen, Xinting Hu, Shuo Wang, Hengheng Zhang, Shaofeng Zhang, Changcheng Li, Boyu Shi, Qi Tian

AI总结:

HERO-MoE通过重用前序层路由分布并采用尺度保持融合机制,在几乎不增加开销的情况下降低MoE训练损失,实验显示最终损失从1.6393降至1.6184。

AI中文摘要:

混合专家(Mixture-of-Experts, MoE)架构已成为在保持计算稀疏的同时扩展模型容量的标准方式,然而路由仍然是决定MoE质量和训练行为的关键因素。先前的实证研究表明,MoE路由反映了跨深度的输入语义和上游计算,但标准路由器并未显式利用前序层产生的路由分布。我们提出HERO-MoE(Historical Expert ROuting with Scale-Preserving Fusion),一种路由框架,通过重用从前序MoE层收集的分离(detached)、密集路由分布,将历史路由先验注入MoE路由器。核心思想很简单:HERO-MoE保留原始的条件于令牌(token-conditioned)的路由分支,并在标准softmax和top-$k$分派之前添加一个残差历史路由贡献。为了稳定这一历史信号,HERO-MoE引入了一种尺度保持融合机制,该机制将历史路由记忆的幅度与当前隐藏表示相匹配,并考虑可见历史层的数量,而无需引入辅助路由损失或融合特定的调优参数。通过重用前序MoE层已计算的路由分布,HERO-MoE以适度的端到端开销改善了训练损失的降低。所得路由器仍与标准稀疏分派(包括top-$k$和组限制路由)兼容,并且可以以最小的架构改动插入现有MoE骨干中。在一个约8B参数、0.5B激活参数的MoE模型上进行的实验(该模型从零开始在100B令牌上训练)表明,HERO-MoE将最终损失从1.6393降至1.6184,而峰值内存和FLOPs仅分别增加0.44%和0.64%。

英文摘要:

Mixture-of-Experts (MoE) architectures have become a standard way to scale model capacity while keeping computation sparse, yet routing remains a key determinant of MoE quality and training behavior. Prior empirical studies suggest that MoE routing reflects input semantics and upstream computation across depth, but standard routers do not explicitly use the routing distributions produced by preceding layers. We propose HERO-MoE, Historical Expert ROuting with Scale-Preserving Fusion, a routing framework that injects historical routing priors into MoE routers by reusing detached, dense routing distributions collected from preceding MoE layers. The key idea is simple: HERO-MoE preserves the original token-conditioned routing branch and adds a residual historical routing contribution before the standard softmax and top-$k$ dispatch. To stabilize this historical signal, HERO-MoE introduces a scale-preserving fusion mechanism that matches the magnitude of historical routing memory to the current hidden representation and accounts for the number of visible historical layers, without introducing an auxiliary routing loss or a fusion-specific tuning parameter. By reusing routing distributions already computed by preceding MoE layers, HERO-MoE improves training-loss reduction with modest end-to-end overhead. The resulting router remains compatible with standard sparse dispatch, including top-$k$ and group-limited routing, and can be inserted into existing MoE backbones with minimal architectural changes. Experiments on an approximately 8B-parameter MoE model with 0.5B active parameters, trained from scratch on 100B tokens, show that HERO-MoE reduces the final loss from 1.6393 to 1.6184, while peak memory and FLOPs increase by only 0.44\% and 0.64\%, respectively.

↑