arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EdgeXpert:一种基于混合专家与推测解码的内存高效型大语言模型推理边缘设备

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding

Sangwoo Ha, Hyunwoo Seo, Yurim Jo, Youngjin Moon, Hoi-Jun Yoo

arXiv 2608.05303首次发表:更新:

发表机构

KAIST(韩国科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

EdgeXpert是一款软硬件协同设计的LLM加速器,通过优化预填充与解码阶段的专家处理,解决推测解码与MoE的兼容性问题,在降低延迟与能耗的同时维持精度,适用于边缘设备的LLM推理。

AI 中文摘要

大语言模型(LLM)的设备端部署已成为个性化边缘应用的核心需求,其主要瓶颈在于前馈网络(FFN)层的外部内存访问(EMA)。推测解码与混合专家(MoE)是极具潜力的解决方案:推测解码通过每阶段生成多个令牌减少解码阶段数,MoE则通过稀疏专家激活降低每阶段成本,但二者结合存在不兼容性。本文提出EdgeXpert,一款软硬件协同设计的LLM加速器,以解决该不兼容问题。在预填充阶段,提示级专家复用将路由重构为提示级专家复用,而非独立的逐令牌专家选择;其通过轻量编码器识别重要令牌,从中构建共享专家集,并以缩减后的专家预算路由较不重要令牌,以降低专家EMA。在解码阶段,深度感知专家合并利用相同深度候选令牌的上下文相似性与互斥性,EdgeXpert仅加载关键通道而非所有所需通道的并集,并应用计算校准以恢复精度,且无需额外内存访问。EdgeXpert采用三星28nm工艺在800 MHz下综合实现,与现有工作相比,其延迟降低达56.3%,能耗降低达44.1%,同时保持接近基线的精度。

英文摘要

On-device deployment of Large Language Models (LLMs) has become essential for personalized edge applications. A primary bottleneck is external memory access (EMA) in feed-forward network (FFN) layers. Speculative decoding and mixture-of-experts (MoE) are promising solutions. Speculative decoding reduces the number of decoding stages by generating multiple tokens per stage, and MoE minimizes per-stage cost through sparse expert activation. However, there is an incompatibility when combining these two techniques. We propose EdgeXpert, a software-hardware co-designed LLM accelerator that resolves this incompatibility. In the prefill stage, the prompt-wise expert reuse reformulates routing as prompt-level expert reuse rather than independent per-token expert selection. It identifies important tokens using a lightweight encoder, constructs a shared expert set from them, and routes less important tokens with a reduced expert budget to lower expert EMA. In the decode stage, depth-aware expert coalescing exploits the contextual similarity and mutual exclusivity of same-depth candidate tokens. Rather than loading the union of all required channels, EdgeXpert loads only salient channels and applies computational calibration to recover accuracy without additional memory access. Synthesized in Samsung 28nm technology at 800 MHz, EdgeXpert achieves up to 56.3% latency reduction and 44.1% energy reduction compared to prior works, while maintaining near-baseline accuracy.

CommentsAccepted at the 59th IEEE/ACM International Symposium on Microarchitecture (MICRO 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑