arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MoE-CORE:面向内存受限MoE推理的协调专家卸载与驻留

MoE-CORE: Coordinated Expert Offloading and Residency for Memory-Constrained MoE Inference

Ke Yang, Yongji Gao, Xushi Li, Kui Luo, Sicheng Zhang, Tianming Zhou, Keyi Liu, Shufang Lu, Aoxuan Chen, Jie Meng, Jingchun Gao, Dan Li, Xinkai You, Dan Li, Zhixiang Xia, Yan Shi, Yang Liu, Yanjia Zeng, Liangjun Feng

arXiv 2610.01950首次发表:更新:

发表机构

University of Chinese Academy of Sciences; Huawei Technologies Co. Ltd; Huawei Ascend Computing Technology Development; Shanghai Jiao Tong University(中国科学院大学; 华为技术有限公司; 华为昇腾计算技术发展中心; 上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MoE-CORE通过协调专家卸载与驻留,在内存受限设备上实现高效的MoE推理,显著降低TPOT,支持精确执行与可选替换路径。

AI 中文摘要

稀疏专家激活降低了MoE模型的计算量,但专家权重可能超出有限的设备内存。卸载使得推理在紧凑型AI设备上可行,但将主机到设备的传输暴露在推理路径中。我们提出MoE-CORE,一个在内存受限的MoE推理中协调专家卸载与驻留的系统。它在预填充阶段将完整的专家层交替存储在缓冲区中。在解码阶段,它结合了非均匀的层间缓存容量、领域信息初始化、路由历史感知的替换以及跨层预取。主配置精确执行路由器选择的专家;一个可选的基于分数的替换路径处理符合条件的低分未命中。主要比较分别使用1K和128个令牌的输出上限用于MoE-CORE和vLLM Prefetch。在DeepSeek-V4-Flash-W4A8上,每个模型五个工作负载,MoE-CORE的平均每输出令牌时间(TPOT)为38.0-44.8毫秒,而评估的vLLM Prefetch配置为1268.9-1269.1毫秒;在GLM-5.2-W4A8C8上,对应值分别为206.6-220.5和5941.5-5941.8毫秒。在84GB NPU内存上限下,最佳测量的DeepSeek GSM8K配置通过近似专家替换和深度2的多令牌预测(MTP)实现了21.5毫秒的TPOT。这些结果支持在设备内存约束下协调专家驻留和传输调度。代码在此。

英文摘要

Sparse expert activation reduces MoE models' computation, yet expert weights can exceed limited device memory. Offloading makes inference feasible on a compact AI appliance but exposes host-to-device transfers to the inference path. We present MoE-CORE, a system that coordinates expert offloading and residency for memory-constrained MoE inference. It stages complete expert layers in alternating buffers during prefill. During decode, it combines nonuniform layer-wise cache capacity, domain-informed initialization, routing-history-aware replacement, and cross-layer prefetching. The main configuration executes router-selected experts exactly; an optional score-based substitution path handles eligible low-score misses. The main comparison uses 1K- and 128-token output caps for MoE-CORE and vLLM Prefetch, respectively. Across five workloads per model, MoE-CORE records a mean time per output token (TPOT) of 38.0-44.8 ms versus 1268.9-1269.1 ms for the evaluated vLLM Prefetch configuration on DeepSeek-V4-Flash-W4A8; the corresponding values on GLM-5.2-W4A8C8 are 206.6-220.5 and 5941.5-5941.8 ms. Under an 84-GB NPU-memory cap, the best measured DeepSeek GSM8K configuration achieves a TPOT of 21.5 ms with approximate expert substitution and multi-token prediction (MTP) at depth 2. These results support coordinated expert residency and transfer scheduling under a device-memory constraint. The code is here.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑