arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34117cs.LGcs.AI

SlimWise:解耦预填充与解码阶段的专家剪枝以实现高效MoE服务

SlimWise: Decoupling Expert Pruning Across Prefill and Decode for Efficient MoE Serving

  • a2sys
  • KAIST(韩国科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

Gunho Park, Kyoungho Jeun, Juntaek Oh, Byeongjun Shin, Baeseong Park, Minsoo Rhu

AI总结:

SlimWise通过为预填充和解码阶段分别定制专家池,利用KV缓存交接和低成本蒸馏,在50%剪枝下将MoE解码吞吐量提升1.81倍,同时保持精度。

AI中文摘要:

混合专家(MoE)模型每个token仅激活少数专家,但批处理解码可能访问几乎整个专家池,使得专家权重流量成为主要瓶颈。专家剪枝可减少该流量,但传统方法也会剪除计算密集的预填充阶段,牺牲模型质量却几乎不提升吞吐量。我们提出SlimWise,一个为每个推理阶段定制专家池的服务框架。SlimWise使用完整模型进行预填充,并使用剪枝模型进行解码,该剪枝模型直接重用预填充生成的KV缓存,无需转换。在两个MoE骨干网络和三种剪枝标准下,这种免训练的KV缓存交接在许多设置中显著缩小了与完整模型相比的精度差距。我们还表明,基准精度可能掩盖剪枝引起的生成长度的显著变化。为解决这些失真和残余精度损失,SlimWise引入了一个低成本的蒸馏阶段,训练解码器从完整模型的KV缓存继续生成,同时仅更新一小部分参数。在vLLM中实现,SlimWise支持预填充-解码(PD)分离和PD共置服务。在Qwen3.6-35B-A3B上,SlimWise在50%专家剪枝下将解码吞吐量提升高达1.81倍,且精度损失极小。

英文摘要:

Mixture-of-experts (MoE) models activate few experts per token, yet batched decoding can access nearly the entire expert pool, making expert-weight traffic a major bottleneck. Expert pruning reduces this traffic, but conventional approaches also prune compute-bound prefill, sacrificing model quality for little throughput benefit. We present SlimWise, a serving framework that tailors the expert pool to each inference phase. SlimWise performs prefill with the full model and decode with a pruned model that directly reuses the prefill-generated KV cache without conversion. Across two MoE backbones and three pruning criteria, this training-free KV cache handoff substantially narrows accuracy gaps relative to the full model in many settings. We also show that benchmark accuracy can conceal substantial pruning-induced changes in generation length. To address these distortions and residual accuracy loss, SlimWise introduces a low-cost distillation stage that trains the decoder to continue from full-model KV caches while updating only a small subset of parameters. Implemented in vLLM, SlimWise supports both prefill-decode (PD) disaggregation and PD-colocated serving. On Qwen3.6-35B-A3B, SlimWise improves decode throughput by up to 1.81x at 50% expert pruning with minimal accuracy loss.

↑