发表机构
StepOs; ShanghaiTech University; Xiamen University; Chongqing University(未知; 上海科技大学; 厦门大学; 重庆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对稀疏MoE模型在有限内存下部署及专家卸载的传输瓶颈问题,提出SpecPrefetch框架,通过分离传输预测与执行路由,结合窗口感知调度器,提升专家召回率和设备解码吞吐量。
AI 中文摘要
稀疏混合专家(MoE)模型通过条件专家激活扩展基础模型容量,但在有限的加速器内存下难以部署完整的专家池。专家卸载虽能缓解内存压力,但会引入依赖路由的传输瓶颈。为此提出SpecPrefetch,它使用共享轻量级适配器仅为异步传输预测下一层专家候选者,冻结的原生路由器仍决定最终执行的专家。通过分离传输预测与执行路由,减少了暴露的专家加载延迟。窗口感知调度器在缓存和带宽约束下对可行传输进行优先级排序。在多个模型基准设置中表现出色,在骁龙8精英设备上还提高了解码吞吐量。
英文摘要
Sparse Mixture-of-Experts (MoE) models expand foundation model capacity through conditional expert activation, but their full expert pools remain difficult to deploy under limited accelerator memory. Although expert offloading alleviates memory pressure by moving inactive experts to host memory or storage, it introduces a routing-dependent transfer bottleneck: required experts are known only after native top-\(K\) routing, which serializes routing, expert loading, and expert execution during inference. To address this bottleneck, we propose SpecPrefetch, a parameter-efficient prefetching framework for offloaded MoE inference. SpecPrefetch uses a shared lightweight adapter to predict next-layer expert candidates only for asynchronous transfer, while the frozen native router still determines the final executed experts. By separating transfer prediction from execution routing, SpecPrefetch reduces exposed expert-loading latency without changing pretrained routing semantics, so prediction errors affect transfer efficiency rather than model outputs. In addition, a window-aware scheduler prioritizes feasible transfers under cache and bandwidth constraints. Across Qwen3-VL-30B-A3B and DeepSeek-VL2-Tiny, SpecPrefetch achieves the best average expert recall in 9 out of 10 model-benchmark settings with substantially fewer trainable parameters than learned predictor baselines. On a Snapdragon 8 Elite device, SpecPrefetch further improves decoding throughput by up to \(20\%\) over a compute-optimized offloading runtime, demonstrating practical benefits for storage-constrained MoE deployment. The code and model weights are available at https://github.com/wei390/SpecPrefetch.