发表机构
Tencent(腾讯)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RAZOR是一种无需训练的专家剪枝方法,通过共识残差评估专家可替换性,在多个大语言模型上以固定预算剪枝时优于现有方法,但生成稳定性仍需关注。
AI 中文摘要
混合专家(MoE)模型每个词元仅激活少数专家,但存储完整的专家池。专家剪枝可减轻这种存储负担;在固定的剪枝预算下,目标是尽可能保持原始模型的输出分布。然而,专家的使用频率或贡献大小本身并不能决定其被移除后造成的损害。关键在于存活的专家计算能否替代其功能。我们提出RAZOR,一种无需训练的专家剪枝方法,利用共识残差(即专家输出与原始加权混合的偏差)来对功能可替换性进行评分。在固定层输入下,精确的单次删除恒等式考虑了幸存者的重新归一化和路由器选择的补充,从而在无需梯度或恢复训练的情况下,通过校准词元聚合的局部评分实现预算化剪枝。在GLM-4.7-Flash、Qwen3.6-35B-A3B、DeepSeek-V4-Flash-0731和Hy3上,在25%和50%的专家移除比例下,RAZOR在所有八个设置中均取得了最高的九任务宏平均分。在两个与REAP基准运行匹配的骨干模型上,RAZOR超过REAP 2.12至5.59个百分点,并赢得了所有36项配对任务比较。在所有四个匹配的GLM-4.7-Flash和Qwen3.6-35B-A3B模型-预算设置中,RAZOR相对于REAP降低了反向KL。然而,对Qwen3.6-35B-A3B生成响应的分析揭示了多样性、格式和终止方面的变化,这强调了任务保留和预测保真度并不能保证生成稳定性。
英文摘要
Mixture-of-experts (MoE) models activate only a few experts per token but store the entire expert pool. Pruning this pool requires identifying experts whose removal preserves model behavior. Routing frequency and output magnitude do not fully describe deletion damage, which also depends on how the surviving and replacement experts compensate for the removed output. We introduce RAZOR, a training-free pruning method based on consensus residuals, the deviations of expert outputs from their original weighted mixture. At a fixed layer input, these residuals give the exact output change for a single deletion under survivor renormalization and router refill. RAZOR aggregates this damage by conditional root mean square and selects experts under a layerwise budget using forward computation alone, without gradients, subset search, or recovery training. Against frequency, activation-norm, and REAP baselines on GLM-4.7-Flash and Qwen3.6-35B-A3B at 25% and 50% expert removal, it attains the highest macro average over nine reasoning-intensive tasks in all four model-budget settings, gaining 2.12-5.59 points over REAP and lowering reverse KL in all four. On DeepSeek-V4-Flash-0731 and Hy3, it also achieves the highest macro average among the three residual criteria. Local exactness does not guarantee better joint pruning. Generation analyses show changes in diversity, formatting, and termination despite higher task scores.