arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.00367cs.LG

MoRA:通过路由器偏置学习与专家近似进行MoE剪枝

MoRA: MoE Pruning via Router Bias Learning and Expert Approximation

Yushuai Sun, Zikun Zhou, Lin Gao, Jun Yu, Wenjie Pei

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出MoRA框架,通过可学习路由器偏置和专家近似机制进行结构化MoE剪枝,在多个模型上移除25%和50%专家,并在九个零样本基准上超越现有剪枝算法。

中文摘要 AI 辅助

混合专家(MoE)模型通过为每个令牌仅激活一小部分专家,实现了在有限每令牌计算量下的参数扩展,但部署这些模型仍需要将完整的专家池加载到内存中。结构化专家剪枝通过移除专家可以有效减少内存占用。然而,现有的剪枝方法要么使用与模型性能不太一致的专家排序标准,要么依赖计算成本高昂的有效专家子集搜索。此外,这些方法通常忽略了保留专家之间路由行为的冗余性。在本文中,我们提出了通过路由器偏置学习与专家近似进行MoE剪枝(MoRA),一个用于结构化MoE专家剪枝的框架。我们为每个专家引入一个可学习的路由器偏置,并通过最小化语言建模损失和路由多样性正则化器来优化这些偏置。学习到的路由器偏置使路由概率分布变得尖锐,以识别对模型性能至关重要的专家,同时鼓励选择具有多样路由偏好的专家。此外,我们引入了一种专家近似机制作为剪枝后的增强。它利用剩余的专家通过仿射变换来近似被剪枝专家的输出,进一步提高剪枝模型的性能。我们在Qwen3-30B-A3B、DeepSeek-V2-Lite和Moonlight-16B-A3B上评估了MoRA,在每个MoE层中移除25%和50%的路由专家。在九个零样本基准上的大量实验表明,MoRA优于最先进的剪枝算法。我们的代码将发布。

英文摘要

Mixture-of-Experts (MoE) models enable parameter scaling with limited per-token computation by activating only a small subset of experts for each token, but deploying them still requires loading the complete expert pool into memory. Structured expert pruning can effectively reduce the memory usage by removing experts. However, existing pruning methods either use expert ranking criteria that are not well aligned with model performance or rely on effective expert subset searching that is computationally expensive. Moreover, these methods typically overlook the routing-behavior redundancy among the retained experts. In this paper, we propose MoE Pruning via Router Bias Learning and Expert Approximation (MoRA), a framework for structured MoE expert pruning. We introduce a learnable router bias for each expert and optimize these biases by minimizing the language-modeling loss and a routing-diversity regularizer. The learned router biases sharpen the routing probability distributions to identify experts critical to model performance while encouraging the selection of experts with diverse routing preferences. In addition, we introduce an expert approximation mechanism as a post-pruning enhancement. It leverages the remaining experts to approximate the outputs of pruned experts by affine transformation, further improving the performance of the pruned model. We evaluate MoRA on Qwen3-30B-A3B, DeepSeek-V2-Lite, and Moonlight-16B-A3B, removing 25\% and 50\% of the routed experts in each MoE layer. Extensive experiments on nine zero-shot benchmarks show that MoRA outperforms state-of-the-art pruning algorithms. Our code will be released.

发表机构

  • Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑