OMP-MoE:基于正交匹配追踪的混合专家大语言模型高效专家剪枝
OMP-MoE: Efficient Expert Pruning for Mixture-of-Experts LLMs via Orthogonal Matching Pursuit
浏览论文内容
中文总结 AI 辅助
提出OMP-MoE,一种无需训练的MoE大模型专家剪枝框架,通过正交匹配追踪重构剪枝问题,结合注水策略优化跨层分配,在25-50%剪枝率下优于现有方法,实现高保真压缩与推理加速。
中文摘要 AI 辅助
混合专家(MoE)模型能够实现大型语言模型的高效扩展,但由于巨大的内存需求,面临关键的部署挑战。现有的剪枝方法要么产生高昂的搜索成本,要么忽略了专家之间的动态相互依赖性。为了解决这些挑战,我们提出了OMP-MoE,一种新颖的无需训练的压缩框架,用于减少基于MoE的大语言模型中的专家冗余。基于对专家贡献模式的观察,我们将剪枝问题重新表述为通过正交匹配追踪解决的稀疏信号重建任务。具体来说,我们的方法首先将单个专家贡献视为字典原子,并以线性计算复杂度贪婪地选择最小化重建误差的专家。然后,我们通过一种考虑重建质量和路由稳定性的注水策略来优化跨层专家分配。最后,我们引入了OMP-MoE†,一种基于能量预测动态调整专家激活的自适应推理机制。在Qwen、DeepSeek-V2、GPT-OSS和Mixtral MoE上的全面实验表明,在25-50%的剪枝比例下,我们的方法始终优于现有方法。对于Qwen3-30B-A3B在50%压缩率下,我们保留了93.3%的原始性能,实现了33倍的搜索加速和1.55倍的推理加速。代码将在论文被接收后公开。
英文摘要
Mixture-of-Experts (MoE) models enable efficient scaling of large language models but face critical deployment challenges due to massive memory requirements. Existing pruning methods either incur prohibitive search costs or neglect the dynamic interdependencies between experts. To address these challenges, we present OMP-MoE, a novel training-free compression framework for reducing expert redundancy in MoE-based LLMs. Based on observations of expert contribution patterns, we reformulate the pruning problem as a sparse signal reconstruction task solved through Orthogonal Matching Pursuit. Specifically, our method first treats individual expert contributions as dictionary atoms and selects experts that greedily minimize reconstruction error with linear computational complexity. Then, we optimize cross-layer expert allocation through a water-filling strategy that accounts for both reconstruction quality and routing stability. Finally, we introduce OMP-MoE†, an adaptive inference mechanism that dynamically adjusts expert activation based on energy prediction. Comprehensive experiments on Qwen, DeepSeek-V2, GPT-OSS, and Mixtral MoE demonstrate consistent improvements over existing methods at 25-50% pruning ratios. For Qwen3-30B-A3B at 50% compression, we retain 93.3% of original performance, achieving 33$\times$ faster search and 1.55$\times$ inference speedup. Codes will be available after acceptance.