arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

无需训练的细粒度混合专家模型中激活专家的减半操作

Training-Free Halving of Activated Experts in Fine-Grained Mixture-of-Experts Models

Xing Chen, Hengshuai Yao

arXiv 2609.04575首次发表:更新:

AI 中文总结

该研究提出无需训练的细粒度 MoE 专家减半方法,通过分离路由与归一化效应,在 Qwen 系列模型上实现专家减半时精度损失极小,为 MoE 压缩提供了有效方案。

AI 中文摘要

现代细粒度混合专家(Mixture-of-Experts, MoE)模型会将每个 token 路由至少量专家,并对其路由器概率进行重新归一化。我们发现,这种重新归一化会隐式地将专家输出增益校准到训练时的 top-$k$:在推理阶段减小 $k$ 不仅会改变所使用的专家,还会改变专家分支的强度。我们通过激活 top $k_1$ 个专家,同时按 top $k_2$ 个专家的概率质量进行归一化,来分离这些效应,该方法仅引入一个整数,无参数、无训练开销且无明显计算开销。在 Qwen3.6-35B-A3B 模型上,采用标准重新归一化时,将专家数量从 8 减至 4 会导致 MMLU 下降 4.65 个点,但当 $k_2=16$ 时仅下降 0.35 个点,同时路由专家的计算量减半。该结果在大 11 倍的 Qwen3.5-397B-A17B 模型上也得到了复现,在使用合适参考集的情况下,将专家数量从 10 减至 5 仅损失 0.55 个点。完全移除重新归一化会造成灾难性后果,表明保留合适的参考质量至关重要。我们进一步发现,困惑度和下游准确率倾向于不同的 $k_2$,警示仅使用未标记文本选择 MoE 压缩设置存在风险。分析还显示,专家身份的重要性远高于专家加权,而平衡且面向领域的路由为专家剪枝留下的空间有限。

英文摘要

Modern fine-grained Mixture-of-Experts (MoE) models route each token to a small number of experts and renormalize their router probabilities. We show that this renormalization implicitly calibrates expert output gain to the training top-$k$: reducing $k$ at inference changes not only which experts are used but also the strength of the expert branch. We separate these effects by activating the top $k_1$ experts while normalizing by the probability mass of the top $k_2$ experts, introducing one integer with no parameters, training, or measurable compute overhead. On Qwen3.6-35B-A3B, reducing from 8 to 4 experts causes a 4.65-point MMLU drop under standard renormalization but only 0.35 points with $k_2=16$, while halving routed-expert compute. The result replicates on the $11\times$ larger Qwen3.5-397B-A17B, where reducing from 10 to 5 experts loses only 0.55 points with an appropriate reference set. Removing renormalization entirely is catastrophic, showing that preserving a suitable reference mass is crucial. We further find that perplexity and downstream accuracy favor different $k_2$, cautioning against selecting MoE compression settings using unlabeled text alone. Analyses also show that expert identity matters substantially more than expert weighting, while balanced and domain-specialized routing leaves limited room for expert pruning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑