arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.16721cs.LG

一半的专家,全部的代码:用于编码的专家混合语言模型的一次性领域剪枝

Half the Experts, All the Code: One-Shot Domain Pruning of Mixture-of-Experts LLMs for Coding

Anik Jha

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对编码任务的专家混合语言模型,通过五种策略对两个模型进行剪枝,探究能移除多少及哪些专家。发现移除一半专家在代码基准测试无损失,损害多在编码外,但获胜策略因模型而异,还揭示了困惑度等问题及剪枝与量化的关系,强调需 per - model 验证。

中文摘要 AI 辅助

最强的开放权重编码模型是专家混合(MoE)网络:其大部分规模来自大量“专家”子网络池,其中只有少数作用于任何令牌。正因如此,这些模型无法在大多数开发者拥有的机器上运行,但对于只想要编码帮助的用户来说,大多数专家编码的能力永远不会被调用。我们研究了通过五种选择策略对来自不同家族的两个近期开放权重MoE模型(Qwen3.6 - 35B - A3B和Gemma - 4 - 26B - A4B)进行剪枝,能移除多少专家以及哪些专家,通过用户判断模型是否仍能编写正确代码的方式来评判。结果表明,从任一模型中移除一半专家,在主要代码基准测试中没有统计学上可检测到的损失,且损害几乎完全落在编码之外的能力上。但获胜策略在两个模型之间会翻转,所以在一个家族上验证有效的方法不能假定在另一个家族上也有效。我们还进一步表明,困惑度这一许多剪枝文献所依赖的指标,可能会将损坏的模型评价得高于完整模型;轻量级微调可以恢复约一半激进剪枝所损失的内容;与将整个模型量化到相同内存相比,剪枝仅在量化必须降至每个权重低于3位时才占优。五次试图推翻这种交叉情况(通过更好的校准、谨慎的选择、因果专家重要性、失败归因以及让每个模型根据执行反馈修复失败的智能体评估)均未成功;最后表明一次性基准测试普遍高估了压缩惩罚,因为一次修复轮次完全消除了2位量化惩罚。专家剪枝有效,但需要针对模型实际服务的任务进行每个模型的验证。

英文摘要

The strongest open-weight coding models are mixture-of-experts (MoE) networks: most of their size comes from large pools of "expert" subnetworks, of which only a few act on any token. That pool is why these models do not fit on the machines most developers own, yet for a user who only wants coding help, most experts encode abilities that will never be invoked. We ask how many experts can be removed, and which, by pruning two recent open-weight MoE models from different families (Qwen3.6-35B-A3B and Gemma-4-26B-A4B) under five selection strategies, judged the way a user would: by whether the model still writes correct code. Half the experts can be removed from either model with no statistically detectable loss on the primary code benchmark, and the damage lands almost entirely on abilities outside coding, the intended trade. But the winning strategy flips between the two models, so a recipe validated on one family cannot be assumed to work on another. We further show that perplexity, the metric much of the pruning literature leans on, can rate a broken model above an intact one; that a lightweight fine-tune recovers about half of what aggressive pruning loses; and that against quantizing the full model to the same memory, pruning wins only where quantization would have to drop below 3 bits per weight. Five attempts to overturn that crossover, with failure criteria fixed in advance (better calibration, guarded selection, causal expert importance, failure attribution, and an agentic evaluation letting each model repair its failures from execution feedback), all leave it standing; the last shows single-shot benchmarks overstate compression penalties broadly, as one repair turn erases the 2-bit quantization penalty entirely. Expert pruning works, but it demands per-model validation on the task the model will actually serve.

补充信息

↑