arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36301cs.LGcs.AIcs.CLstat.ML

MoRE: 通过硬件感知的低秩路由扩展混合专家模型

MoRE: Scaling mixture of experts with hardware-aware low-rank routing

  • University of Pennsylvania(宾夕法尼亚大学)
  • The Wharton School, University of Pennsylvania(宾夕法尼亚大学沃顿商学院)

机构由 AI 辅助整理,请以论文原文为准。

Honam Wong, Surbhi Goel, Enric Boix-Adserà

AI总结:

MoRE通过低秩路由分解降低混合专家模型的路由成本,证明对数秩足以保持表达能力,并在不损害记忆与推理能力的前提下提升专家数量与问答性能。

AI中文摘要:

混合专家(MoE)层是前沿语言模型的核心组成部分,最近的架构趋向于使用更多且更小的专家。在这种机制下,标准的线性路由器成为瓶颈:对于$M$个专家和隐藏维度$h$,其每个token的成本为$\Theta(Mh)$,一旦$M$很大,该成本便主导了MoE层。我们引入了MoRE(秩缩减路由的混合专家模型),它将路由器权重矩阵分解为秩$r$,并将路由成本降低至$O((h + M)r)$。我们证明了当激活专家数量固定时,秩为$M$的对数级别就足以保证路由的表达能力,并且在精度因素范围内这是必要的。我们还证明了在对数秩下,高斯记忆模型中的负载均衡得以保持,并且在合成电话簿任务上的训练表明,低秩不会损害记忆能力。在匹配的激活FLOPs下,该分解允许专家数量增加$\Theta(h/r)$倍。为了在墙钟时间上实现这一增益,我们在推理时设计了一个融合的Triton内核,避免了HBM上昂贵的存储操作。实验上,MoRE在电话簿任务上提升了记忆能力,并在预训练后提升了知识密集型问答基准的性能,同时保持了推理能力。代码可在https://this URL获取。

英文摘要:

Mixture-of-Experts (MoE) layers are central to frontier language models, and recent architectures push toward more and smaller experts. In this regime, the standard linear router becomes a bottleneck: with $M$ experts and hidden dimension $h$, its per-token cost $Θ(Mh)$ dominates the MoE layer once $M$ is large. We introduce MoRE (Mixture of Rank-reduced-routed Experts), which factorizes the router weight matrix at rank $r$ and reduces the routing cost to $O((h + M)r)$. We prove that rank logarithmic in $M$ suffices for routing expressivity when the number of active experts is fixed, and is necessary up to precision factors. We also prove that logarithmic rank preserves load balance in a Gaussian memorization model, and training on a synthetic phonebook task shows that low rank does not hurt memorization. At matched active FLOPs, the factorization allows a factor of $Θ(h/r)$ more experts. To realize this gain in wall-clock time, we design a fused Triton kernel at inference that avoids expensive memory operations on HBM. Empirically, MoRE improves memorization on the phonebook task and performance on knowledge-intensive Q\&A benchmarks after pretraining, while matching reasoning ability. Code available at https://github.com/Matheart/MoRE_code.

↑