arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.24434cs.LGcs.AI

DraftExpert:用于终端设备混合专家推理的扩展感知自推测解码

DraftExpert: Expansion-Aware Self-Speculative Decoding for End-Device MoE Inference

Dengke Han

首次发表
浏览论文内容

中文总结 AI 辅助

针对终端设备MoE推理中自推测解码的新瓶颈,提出DraftExpert框架,通过自蒸馏训练轻量级草稿专家,在不同卸载场景下,该框架有效提高了解码吞吐量、草稿接受率和预取命中率。

中文摘要 AI 辅助

大混合专家(MoE)语言模型对终端设备部署具有吸引力,因其每个token只有一小部分专家活跃,但路由的专家权重常超加速器内存。在延迟关键的单用户设置中,自推测解码面临新瓶颈。我们提出DraftExpert,一种用于专家卸载MoE推理的扩展感知自推测解码框架。它通过自蒸馏来自冻结目标MoE的残差、logit/token和路由器协议信号,每层训练一个轻量级驻留在加速器的草稿专家。实验表明,在不同卸载场景下,DraftExpert平均提高解码吞吐量1.45倍,将草稿接受率提高到84% - 87%,预取命中率达86% - 88%。

英文摘要

Large Mixture-of-Experts (MoE) language models are attractive for end-device deployment because only a small subset of experts is active per token, but their routed expert weights often exceed accelerator memory. We target latency-critical single-user settings where routed experts are staged on demand from CPU memory to a GPU or from Flash to a mobile NPU. In this setting, self-speculative decoding faces a new bottleneck: increasing the draft expert set improves accuracy but triggers extra expert loading, while cheap small-footprint drafts have low acceptance; moreover, verifying a multi-token block activates the union of target experts and is no longer close to one target step. We propose DraftExpert, an expansion-aware self-speculative decoding framework for expert-offloaded MoE inference. DraftExpert trains one lightweight accelerator-resident draft expert per layer by self-distilling residual, logit/token, and router-agreement signals from the frozen target MoE. At inference time, it uses a fixed-footprint shared+top-1+draft-expert drafter together with confidence--expansion truncation and target-expert prefetching, while final tokens are still exactly verified by the target model. On DeepSeek-V2-Lite and Moonlight-16B-A3B across CPU-GPU and Flash-NPU offload, DraftExpert improves decode throughput by 1.45x on average, raises draft acceptance to 84~87%, and achieves 86~88% prefetch hit rates.

补充信息

↑