发表机构
University of Chinese Academy of Sciences; Huawei Technologies Co. Ltd; Huawei Ascend Computing Technology Development; Shanghai Jiao Tong University(中国科学院大学; 华为技术有限公司; 华为昇腾计算技术发展中心; 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MoE-CORE通过协调专家卸载与驻留,在内存受限设备上实现高效的MoE推理,显著降低TPOT,支持精确执行与可选替换路径。
AI 中文摘要
稀疏专家激活降低了MoE模型的计算量,但专家权重可能超出有限的设备内存。卸载使得推理在紧凑型AI设备上可行,但将主机到设备的传输暴露在推理路径中。我们提出MoE-CORE,一个在内存受限的MoE推理中协调专家卸载与驻留的系统。它在预填充阶段将完整的专家层交替存储在缓冲区中。在解码阶段,它结合了非均匀的层间缓存容量、领域信息初始化、路由历史感知的替换以及跨层预取。主配置精确执行路由器选择的专家;一个可选的基于分数的替换路径处理符合条件的低分未命中。主要比较分别使用1K和128个令牌的输出上限用于MoE-CORE和vLLM Prefetch。在DeepSeek-V4-Flash-W4A8上,每个模型五个工作负载,MoE-CORE的平均每输出令牌时间(TPOT)为38.0-44.8毫秒,而评估的vLLM Prefetch配置为1268.9-1269.1毫秒;在GLM-5.2-W4A8C8上,对应值分别为206.6-220.5和5941.5-5941.8毫秒。在84GB NPU内存上限下,最佳测量的DeepSeek GSM8K配置通过近似专家替换和深度2的多令牌预测(MTP)实现了21.5毫秒的TPOT。这些结果支持在设备内存约束下协调专家驻留和传输调度。代码在此。
英文摘要
Sparse expert activation reduces MoE models' computation, yet expert weights can exceed limited device memory. Offloading makes inference feasible on a compact AI appliance but exposes host-to-device transfers to the inference path. We present MoE-CORE, a system that coordinates expert offloading and residency for memory-constrained MoE inference. It stages complete expert layers in alternating buffers during prefill. During decode, it combines nonuniform layer-wise cache capacity, domain-informed initialization, routing-history-aware replacement, and cross-layer prefetching. The main configuration executes router-selected experts exactly; an optional score-based substitution path handles eligible low-score misses. The main comparison uses 1K- and 128-token output caps for MoE-CORE and vLLM Prefetch, respectively. Across five workloads per model, MoE-CORE records a mean time per output token (TPOT) of 38.0-44.8 ms versus 1268.9-1269.1 ms for the evaluated vLLM Prefetch configuration on DeepSeek-V4-Flash-W4A8; the corresponding values on GLM-5.2-W4A8C8 are 206.6-220.5 and 5941.5-5941.8 ms. Under an 84-GB NPU-memory cap, the best measured DeepSeek GSM8K configuration achieves a TPOT of 21.5 ms with approximate expert substitution and multi-token prediction (MTP) at depth 2. These results support coordinated expert residency and transfer scheduling under a device-memory constraint. The code is here.