S2-MoE:在边缘设备上实现混合专家模型的高效自推测解码
S2-MoE: Enabling Efficient Self-Speculative Decoding for Mixture-of-Experts on Edge Devices
- Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院)
- School of Integrated Circuits, Peking University(北京大学集成电路学院)
- School of Electronics Engineering and Computer Science, Peking University(北京大学电子工程与计算机科学学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
S2-MoE是面向边缘设备MoE推理的高效自推测解码框架,通过路由感知自适应扩展、复用感知门控等设计,使MoE推理在边缘设备上最高获5.3倍加速,平均约2.0倍。
AI中文摘要:
在边缘设备上部署大语言模型(LLMs)进行推理面临严峻挑战,原因在于内存与带宽约束极为严苛。尽管已提出推测解码与混合专家模型(Mixture-of-Experts, MoE)以提升推理效率,但将二者直接结合往往会产生过高的验证开销,且专家复用效果不佳,限制了其在内存受限的边缘场景中的有效性。本研究提出S2-MoE,这是一种用于边缘设备上MoE推理的高效自推测解码框架。S2-MoE通过感知路由的自适应推测扩展减少冗余验证,通过感知复用的专家门控提升验证效率,并通过共享上下文对齐草稿与目标执行。该框架在相关实现代码中,在边缘设备上针对多种MoE模型与数据集,相比标准自回归解码实现了最高5.3倍的加速(平均约2.0倍),相关代码可在指定链接获取。
英文摘要:
Deploying large language models (LLMs) for inference on edge devices is challenging due to severe memory and bandwidth constraints. While speculative decoding and Mixture-of-Experts (MoE) have been proposed to improve inference efficiency, naively combining them often incurs excessive verification overhead and poor expert reuse, limiting their effectiveness in memory-bound edge settings. In this work, we propose S2-MoE, an efficient self-speculative decoding framework for MoE inference on edge devices. S2-MoE reduces redundant verification through routing-aware adaptive speculative expansion, improves verification efficiency with reuse-aware expert gating, and aligns draft and target execution via shared context. Implemented in llama$.$cpp, S2-MoE achieves up to $5.3\times$ speedup (about $2.0\times$ on average) over standard autoregressive decoding across diverse MoE models and datasets on edge devices. Code is available at https://github.com/angerybob/S2-MoE.