SPICE:基于低秩专家代理与异构编排的MoE推理加速投机预取框架
SPICE: Speculative Prefetching with Low-Rank Expert Surrogates and Heterogeneous Orchestration for MoE Inference Acceleration
浏览论文内容
中文总结 AI 辅助
SPICE是用于MoE推理加速的投机预取框架,通过轻量级专家预测与异构编排,在DeepSeek-V2-Lite等模型上实现最高3.12倍的TPOT加速,且质量损失极小。
中文摘要 AI 辅助
混合专家(MoE)模型因稀疏激活特性将模型容量与计算成本解耦,正越来越多地被用于大语言模型(LLM)中。然而,专家参数的庞大内存占用往往超出GPU的显存容量,导致推理延迟主要由主机到设备的PCIe专家加载传输所主导。为应对这些挑战,本文提出了SPICE,这是一种用于MoE卸载的投机预取框架,结合了轻量级专家预测与感知置信度的CPU-GPU异构编排。一方面,SPICE构建了与目标MoE架构对齐的轻量级草稿模型,采用感知置信度的自适应前瞻算法预取高置信度专家;另一方面,当投机预测未命中时,SPICE切换至感知成本的CPU-GPU异构编排:低置信度未命中由驻留共享专家通过低秩专家(LoRE)代理进行近似处理,而精确的剩余工作则被卸载至CPU,并与正在进行的GPU计算异步并行执行。在DeepSeek-V2-Lite和Qwen2-57B-A14B上,跨多种GPU平台的评估显示,SPICE在每输出令牌时间(TPOT)上实现了最高3.12倍的加速,且质量损失极小,这表明有效的MoE卸载不仅需要预测未来的专家,还需要决定哪些未命中值得近似、哪些需要精确恢复,以及精确的剩余工作应在何处执行。
英文摘要
Mixture-of-Experts (MoE) models are increasingly used in LLMs because sparse activation decouples model capacity from compute cost. However, the large expert parameter footprint often exceeds GPU memory capacity, making inference latency dominated by the host-to-device PCIe transfers for expert loading. To address these challenges, this paper presents SPICE, a speculative prefetching framework for MoE offloading that combines lightweight expert prediction with confidence-aware CPU-GPU orchestration. On one hand, SPICE builds a lightweight draft model aligned with the target MoE architecture, using a confidence-aware adaptive lookahead algorithm to prefetch high-confidence experts. On the other hand, when speculative predictions miss, SPICE switches to a cost-aware CPU-GPU heterogeneous orchestration: low-confidence misses are approximated by the resident shared expert with low rank expert (LoRE) surrogates, while exact residual work is offloaded to the CPU and executed asynchronously in parallel with ongoing GPU computation. Evaluated on DeepSeek-V2-Lite and Qwen2-57B-A14B across diverse GPU platforms, SPICE achieves up to 3.12 speedup in Time Per Output Token (TPOT) with minimal quality loss, showing that effective MoE offloading requires not only predicting future experts, but also deciding which misses deserve approximation, which require exact recovery, and where exact residual work should execute.