arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25809cs.LGcs.AI

你只需要所选专家的2/3:细粒度MoE大语言模型中动态专家剪枝的实证研究

You Only Need 2/3 of the Chosen Experts: An Empirical Study of Dynamic Expert Pruning in Fine-Grained MoE LLMs

Yuanteng Chen, Qiwei Lai, Chen Tianqi, Peisong Wang, Yuantian Shao, Nanxin Zeng, Zhilei Liu, Chuangyi Li, Jing Liu, Jian Cheng

首次发表
浏览论文内容

中文总结 AI 辅助

本研究实证考察细粒度MoE大模型的动态专家剪枝,发现统一保留约三分之二所选专家即可保持98.8%性能,并明确了动态分配在激进剪枝下的价值及模型敏感性差异。

中文摘要 AI 辅助

细粒度混合专家(MoE)架构已成为开放权重大语言模型的主流设计,拥有数百个专家,且每个词元选择的专家数量日益增多。这一转变使得动态专家剪枝成为降低推理成本的一条有吸引力的途径。然而,现有证据主要来自较粗粒度的架构和基于似然评分的选择题基准,在细粒度场景下留下了三个核心问题:每个词元的专家选择冗余程度如何,现有剪枝方法在多大程度上有效利用这种冗余,以及什么因素决定了模型对剪枝的敏感性。我们通过对跨越九个架构家族的十二个细粒度MoE检查点进行系统性实证研究来填补这一空白,核心基准套件包含十一个基准,涵盖知识问答、数学、代码生成和通用推理。我们发现,专家选择的冗余程度远超领域当前运行点所假设的水平:统一保留所选专家的大约三分之二,平均可保持未剪枝性能的98.8%,仅需改变一个整数参数,并在两个服务后端上实现了1.2-1.7倍的实测加速。这一简单基线在保守预算下几乎没有为动态分配留下空间:即使是最佳已发表规则,在匹配的专家预算下与基线差异也不足1%。它们的价值在激进剪枝下显现,此时最佳规则相比统一截断可恢复高达3.0%的性能,且收益集中在退化最严重的生成任务中。对激进剪枝的敏感性还取决于模型:更大和具有思考能力的模型更具韧性,而多模态模型则更为脆弱。综合来看,这些发现揭示了细粒度MoE可以省去多少专家计算量,并明确了动态分配何时值得其复杂性,为实际部署和未来剪枝方法提供了信息。

英文摘要

Fine-grained mixture-of-experts (MoE) architectures have become a mainstream design for open-weight LLMs, with hundreds of experts and increasingly many selected per token. This shift makes dynamic expert pruning an attractive route to cheaper inference. Yet existing evidence comes largely from coarser architectures and likelihood-scored multiple-choice benchmarks, leaving three central questions open in the fine-grained regime: how redundant per-token expert selection is, how effectively existing pruning methods exploit that redundancy, and what governs a model's sensitivity to pruning. We fill this gap with a systematic empirical study of twelve fine-grained MoE checkpoints spanning nine architecture families, with a core suite of eleven benchmarks covering knowledge QA, mathematics, code generation, and general reasoning. We find that expert selection is far more redundant than the field's operating points assume: uniformly retaining about two thirds of the selected experts preserves 98.8% of unpruned performance on average, requiring only a one-integer change and delivering 1.2-1.7x measured speedup across two serving backends. This simple baseline leaves little room for dynamic allocation at conservative budgets: even the best published rules differ from it by under 1% at matched expert budgets. Their value emerges under aggressive pruning, where the best rules recover up to 3.0% over uniform truncation, with gains concentrated in the generative tasks that suffer the sharpest degradation. Sensitivity to aggressive pruning also depends on the model: larger and thinking models are more resilient, whereas multimodal models are more vulnerable. Together, these findings reveal how much expert computation fine-grained MoEs can dispense with, and establish when dynamic allocation earns its complexity, informing both practical deployment and future pruning methods.

发表机构

  • Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
  • School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
  • Zhongguancun Academy(中关村学院)
  • University of Science and Technology of China(中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑