arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越免重训练的MoE压缩:压缩后调整的成本归一化研究

Beyond Retraining-Free MoE Compression: A Cost-Normalized Study of Post-Compression Adjustment

Sieun Hyeon, Jaeyoung Do

arXiv 2609.06076首次发表:更新:

发表机构

Seoul National University(首尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出压缩后调整观点,通过小数据预算下的微调或KD,在多个MoE模型和基准上验证了其能显著恢复压缩损失,且全参数微调成本效益最佳。

AI 中文摘要

免重训练的MoE压缩通过剪枝或合并专家来减少部署内存,但通常将压缩后的检查点视为最终产物。我们认为这种观点是不完整的:压缩后的MoE检查点应被更好地理解为压缩初始化,其受益于一个微小的压缩后调整阶段。在两个MoE大语言模型骨干、四种剪枝/合并方法、三种专家保留比例和28个基准测试中,我们在匹配的小数据预算和实测GPU成本下比较了LM微调和基于教师的KD。仅使用3,000个C4样本和单轮调整,Full FT平均恢复了原始与压缩性能差距的37.3%。此外,LM微调比标准token级KD更具成本效益,而在测试的范围内,全参数调整提供了最强的成本-恢复权衡。这些结果表明,免重训练的压缩应与小幅压缩后调整相结合,以恢复压缩过程中损失的相当一部分性能。

英文摘要

Retraining-free MoE compression reduces deployment memory by pruning or merging experts, but often treats the compressed checkpoint as the final artifact. We argue that this view is incomplete: compressed MoE checkpoints are better understood as compressed initializations that benefit from a tiny post-compression adjustment stage. Across two MoE LLM backbones, four pruning/merging methods, three expert-retention ratios, and 28 benchmarks, we compare LM fine-tuning and teacher-based KD under matched small-data budgets and measured GPU costs. Using only 3,000 C4 examples and a single epoch of adjustment, Full FT recovers 37.3% of the original-to-compressed performance gap on average. Moreover, LM fine-tuning is more cost-effective than standard token-level KD, and full-parameter adjustment gives the strongest cost--recovery trade-off among the tested scopes. These results suggest that retraining-free compression should be paired with small post-compression adjustment to recover a substantial portion of the performance lost during compression.

CommentsAccepted to EMNLP 2026 (Main Conference)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑