轻量级微调下的路由器敏感度可识别混合专家模型中可剪枝的专家
Router Sensitivity Under Lightweight Fine-Tuning Identifies Prunable Experts in Mixture-of-Experts Models
浏览论文内容
中文总结 AI 辅助
本研究提出利用轻量级微调下MoE模型的路由器敏感度识别可剪枝专家,通过参数高效适配器微调后按路由器ℓ₂变化排序剪枝,在多个模型和基准上实现低精度损失的高压缩率,使理论驱动的专家剪枝可大规模落地。
中文摘要 AI 辅助
混合专家(Mixture-of-Experts, MoE)模型将总参数量与每token计算量解耦,但部署时仍需存储所有专家。近期理论表明,剪枝微调期间路由器范数变化最小的专家可保留精度,但该结论假设采用全量微调。我们测试了轻量级适配能否恢复这一信号。我们使用参数高效适配器进行简短微调,根据诱导的ℓ₂路由器变化对专家排序,并一次性剪枝变化最小的专家。在Mixtral-8×7B-Instruct(MMLU-Pro精度44.83%)上,仅路由器的LoRA仅训练0.002%的参数,在移除一半专家的情况下,精度优于同秩的全模块LoRA(27.54% vs 24.42%);随着适配扩展到注意力和专家权重,信号质量会下降。精度随LoRA秩单调提升,最高达28.76%。保持路由器权重冻结的IA3方法效果与直接路由器适配相当,而无约束的加性适配器会降低信号质量。路由器引导的MMLU-Pro精度呈准线性下降而非骤降,在最大压缩率下仍约为基于幅度剪枝或随机剪枝的1.8倍,同时减少49%的内存和37%的每token延迟。在25%压缩率下,精度保留率与使用全激活统计量的方法相当。该准则还可迁移到针对数学微调的Qwen1.5-MoE,在移除一半专家的情况下,在11个基准上保留49.7%的平均精度,而随机剪枝的精度降至个位数。因此,轻量级微调下的路由器敏感度使具有理论动机的专家剪枝在大规模场景下具备实用性。
英文摘要
Mixture-of-Experts (MoE) models decouple parameter count from per-token compute, but deployment still requires hosting every expert in memory. Recent theory shows that experts whose router weights change least during fine-tuning can be pruned with provable accuracy preservation, yet the guarantee assumes full fine-tuning. We show that the signal can be elicited through a brief parameter-efficient adaptation. We fine-tune with a lightweight adapter, rank experts by the induced router change, and prune the least-changed experts in one shot. On Mixtral-8$\times$7B-Instruct, router-only LoRA trains 0.002% of parameters and retains 27.54% MMLU-Pro accuracy with half the experts removed, against roughly 16% for magnitude and random pruning. Signal quality improves monotonically with adapter size, reaching 28.76%, and declines as adaptation spreads beyond the router. Under their shared budget, IA3 reaches 28.04% while Houlsby reaches 25.39%. The criterion transfers to Qwen1.5-MoE fine-tuned for mathematical reasoning, retaining 49.7% mean accuracy over eleven benchmarks with half the experts removed. Structural pruning reduces memory by 49% and per-token latency by 37%. Lightweight router sensitivity therefore makes provably motivated, task-conditioned expert pruning practical at scale.
发表机构
- Columbia University(哥伦比亚大学)
- Data Science Institute, Columbia University(哥伦比亚大学数据科学研究所)
- Department of Computer Science, Columbia University(哥伦比亚大学计算机科学系)
机构由 AI 辅助整理,请以论文原文为准。