arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11688cs.ARcs.AIcs.LG

APEX:面向内存高效型边缘MoE推理的自适应专家预取技术

APEX: Adaptive Expert Prefetching for Memory-Efficient Edge MoE Inference

Alish Kanani, Layan Badawi, Umit Y. Ogras

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出APEX自适应专家预取框架,通过预测候选专家动态预取,在保证或可忽略精度损失下,降低边缘MoE推理延迟、提升能效。

中文摘要 AI 辅助

混合专家(Mixture-of-Experts,MoE)模型因在每个token仅激活小部分参数即可提供高模型容量,提升了计算效率,而在边缘部署中颇具吸引力。然而,边缘端的MoE推理受内存限制存在根本性瓶颈:专家参数规模庞大,受容量、成本和功耗约束常驻留在片外内存,导致专家加载成为关键路径。本文提出APEX:自适应专家预取(Adaptive Expert Prefetching),这是一种将专家加载与有效计算重叠的预测资源管理框架。APEX引入轻量型预取路由器,在注意力模块前预测候选专家,通过学习到的置信度模型动态获取额外专家。该自适应策略实现了99%以上的重叠准确率,显著优于固定top-k预取技术。APEX支持两种执行模式:一是保留正确性的模式,保证精确的路由语义;二是无停顿模式,通过对可用专家操作消除剩余停顿,对应用准确率影响可忽略不计。在多个MoE模型上,保留正确性模式相比现有最优基线,将每个token的延迟降低最多26%,能量延迟乘积(EDP)提升最多41%;无停顿模式则进一步提升效率,且对应用准确率影响可忽略不计。这些结果表明,自适应、置信度驱动的专家预取是边缘系统高效MoE推理的有效方法。

英文摘要

Mixture-of-Experts (MoE) models are attractive for edge deployment because they provide high model capacity while activating only a small subset of parameters per token, improving compute efficiency. However, MoE inference at the edge is fundamentally limited by memory. Expert parameters are large and often reside in off-chip memory due to capacity, cost, and power constraints, putting expert loading to the critical path. We present APEX: Adaptive Expert Prefetching, a predictive resource management framework that overlaps expert loading with useful computation. APEX introduces a lightweight prefetch router that predicts candidate experts before the attention block to dynamically fetch additional experts using a learned confidence model. This adaptive strategy achieves over 99% overlap accuracy, significantly outperforming fixed top-k prefetching techniques. APEX supports two execution modes: a correctness-preserving mode that guarantees exact routing semantics, and a stall-free mode that eliminates residual stalls by operating on available experts with negligible impact on application accuracy. Across multiple MoE models, the correctness-preserving mode reduces per-token latency by up to 26% and improves energy-delay product (EDP) by up to 41% over state-of-the-art baselines, while the stall-free mode provides additional efficiency gains with negligible impact on application accuracy. These results establish adaptive, confidence-driven expert prefetching as an effective approach for efficient MoE inference on edge systems.

发表机构

  • University of Wisconsin–Madison(威斯康星大学麦迪逊分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑