混合专家语言模型中专家高阶剪枝
Higher-order pruning of experts in mixture-of-experts language models
- AI Fundamental Research(AI基础研究)
- AWS Agentic AI(AWS智能体AI)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对MoE语言模型参数冗余,提出二阶剪枝方法HOPE,最小化剪枝误差上界,在高达122B参数的模型上优于现有方法,尤其在50%剪枝率下平均排名领先,实现高效压缩。
AI中文摘要:
混合专家(MoE)语言模型存在参数量过大的问题,这造成了显著的内存瓶颈。专家剪枝是减少参数量的最直接方法,然而现有方法独立地对每个专家做出剪枝决策,并假设专家的贡献是纯粹可加的。实际上,MoE中的专家使用本质上是协作性的。我们提出了HOPE(专家高阶剪枝),一个二阶剪枝目标,它可证明地最小化剪枝导致误差的上界。我们证明REAP(一种最先进的一阶剪枝方法)是HOPE在忽略交互项时的特例。在三个前沿MoE模型(参数量高达122B)、两个不同的校准集以及多个基准(包括数学、指令遵循、编码和智能体套件)上,我们证明HOPE比现有方法产生更好的剪枝决策,并且其优势在高剪枝率和具有挑战性的智能体工作负载上最为显著。在50%剪枝率下,HOPE优于所有基线,并在5种方法中取得平均排名1.58(而次优方法REAP为2.42),在智能体编码上增益高达+6.1%。在所有条件下,HOPE再次取得最佳平均排名,并在大多数两两比较中超越所有其他方法。通过保留一阶方法忽略的协作性专家结构,HOPE能够在最小性能下降的情况下实现激进压缩,特别是在涉及长序列中多种专家组合的复杂任务上。
英文摘要:
Mixture-of-Experts (MoE) language models suffer from large parameter counts, which create a significant memory bottleneck. Expert pruning is the most direct approach for reducing this parameter count, yet existing methods make pruning decisions for each expert independently, and assume experts' contributions are purely additive. In reality, expert usage in MoEs is inherently cooperative. We derive HOPE (Higher-Order Pruning of Experts), a second-order pruning objective which provably minimizes an upper bound on the error resulting from pruning. We show that REAP (a state-of-the-art first-order pruning method) is a special case of HOPE where interaction terms are ignored. Across three frontier MoE models (up to 122B parameters), two distinct calibration sets, and multiple benchmarks (including math, instruction following, coding, and an agentic suite), we demonstrate that HOPE produces better pruning decisions than existing methods, and its advantage is most pronounced at high pruning rates and on challenging agentic workloads. At 50% pruning, HOPE outperforms all baselines and achieves an average rank of 1.58 out of 5 methods (versus 2.42 for the next-best method, REAP), with gains of up to +6.1% on agentic coding. Over all conditions, HOPE again achieves the best average rank and surpasses every other method in the majority of head-to-head comparisons. By preserving cooperative expert structure that first-order methods ignore, HOPE enables aggressive compression with minimal degradation, particularly on complex tasks where diverse expert combinations are invoked over long sequences.