发表机构
Capital One(第一资本金融公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对过分散路由下 MoE 模型专家剪枝的标准假设失效问题,提出 MESA 方法,可最小化最坏情况领域性能下降,在多基准任务上优于基线且可推广至多种模型。
AI 中文摘要
专家剪枝通过移除路由识别出的低重要性专家,降低混合专家(Mixture-of-Experts,MoE)模型的内存与推理成本,该方法假设路由概率能提供可靠的重要性信号。我们发现,在训练期间采用激进负载均衡导致的过分散路由场景中,这一假设不再成立:此时 token 几乎均匀分布在所有专家间,重要性信号崩溃。在该场景下,困惑度无法预测下游任务准确率:在 gpt-oss-20B 模型上,困惑度最低的剪枝配置会导致数学推理性能最差,而困惑度最高的配置却能保留该能力;而在标准路由(如 Mixtral-8x7B-Instruct)下,困惑度与准确率会同步下降。过分散路由下的剪枝还会暴露能力权衡问题,即没有单一评分指标能占据优势:激活感知评分可保留数学推理,但会严重降低知识密集型科学任务性能(在 GPQA 上存在 18 分的差距),而基于频率的评分则呈现相反结果。我们提出极小极大专家评分分配(Minimax Expert Score Allocation,MESA),这是一种领域感知方法,会迭代提升受影响最严重领域对应的专家的重要性评分,以最小化最坏情况下的领域性能下降,而非平均准确率。在剪枝 25% 专家时,MESA 实现了各领域最小的最坏情况下降,在 11 个基准任务中,其性能优于激活感知基线的任务达 7 个,同时内存占用相应降低,且该方法可推广至 gpt-oss-120B、Gemma-4-26B-A4B 和 OLMoE-1B-7B 模型。我们的结果表明,过分散路由是一种性质截然不同的剪枝场景,其中标准假设失效,识别该场景是对负载均衡 MoE 模型进行合理专家剪枝的前提条件。
英文摘要
Expert pruning reduces the memory and serving cost of Mixture-of-Experts (MoE) models by removing low-importance experts identified by the router, assuming router probabilities provide a reliable importance signal. We observe that this assumption breaks down under over-dispersed routing, a regime associated with aggressive load-balancing during training, in which tokens are distributed nearly uniformly across experts and importance signals collapse. In this regime, perplexity does not predict downstream task accuracy: on gpt-oss-20B, the lowest-perplexity pruning configuration yields the worst mathematical reasoning, while the highest-perplexity configuration preserves it. This does not occur under standard routing (e.g., Mixtral-8x7B-Instruct), where perplexity and accuracy degrade together. Pruning under over-dispersed routing also exposes a capability trade-off in which no single scoring metric dominates: activation-aware scoring preserves mathematical reasoning but severely degrades knowledge-intensive science (an 18-point gap on GPQA), whereas frequency-based scoring exhibits the reverse. We propose Minimax Expert Score Allocation (MESA), a domain-aware method that iteratively boosts importance scores for experts serving whichever domain is currently worst-affected, minimizing worst-case domain degradation rather than average accuracy. At 25% expert pruning MESA achieves the smallest worst-case degradation across domains, outperforming activation-aware baselines on 7 of 11 benchmarks at a correspondingly reduced memory footprint, and it generalizes to gpt-oss-120B, Gemma-4-26B-A4B, and OLMoE-1B-7B. Our results indicate that over-dispersed routing is a qualitatively distinct pruning regime in which standard assumptions fail, and that recognizing it is a prerequisite for principled expert pruning of load-balanced MoE models.
Comments22 pages, 7 figures. Preprint