发表机构
College of Computer Science and Technology, National University of Defense Technology(国防科技大学计算机学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
HiNa-MoE提出非侵入式CPU算子库,利用Intel AMX优化MoE推理,通过微内核、NUMA感知划分和矩阵运算转换,实现FFN内核3.37倍和端到端2.09倍加速。
AI 中文摘要
混合专家(MoE)推理越来越多地部署在本地和内部环境中,在这些环境中,专家参数往往超过GPU内存容量。在延迟敏感、低并发场景下,反复将路由专家权重从CPU内存暂存到GPU可能代价高昂,使得路由专家前馈网络(FFN)留在多插槽CPU的关键路径上。现有的CPU加速通常依赖于侵入性的、硬件或拓扑特定的要求,例如AMX特定的权重布局或手动NUMA感知放置。这些要求降低了可移植性,并使与标准CPU-GPU卸载管道的集成复杂化。我们提出了HiNa-MoE,一个用于在具有Intel AMX的CPU上进行MoE推理的高性能、非侵入式算子库。HiNa-MoE(1)利用AMX,通过优化的微内核将专家权重保持在标准布局中,而将轻量级布局变换融合到令牌收集和存储中;(2)在简单的页交错策略下应用NUMA感知任务划分,而不修改框架分配器;(3)将解码阶段的矩阵-向量操作转换为小矩阵-矩阵执行,以利用AMX。在多个MoE模型上,HiNa-MoE在FFN内核上实现了高达3.37倍的加速,端到端推理速度比最先进的基线提高了2.09倍,同时与现有框架和部署工作流保持即插即用。
英文摘要
Mixture-of-Experts (MoE) inference is increasingly deployed in local and on-premise environments, where expert parameters often exceed GPU memory capacity. In latency-sensitive, low-concurrency settings, repeatedly staging routed-expert weights from CPU memory to the GPU can be prohibitive, leaving routed-expert feed-forward networks (FFNs) on the critical path of multi-socket CPUs. Existing CPU accelerations often rely on intrusive, hardware- or topology-specific requirements, such as AMX-specific weight layouts or manual NUMA-aware placement. These requirements reduce portability and complicate integration with standard CPU-GPU offloading pipelines. We present HiNa-MoE, a high-performance, non-intrusive operator library for MoE inference on CPUs with Intel AMX. HiNa-MoE (1) exploits AMX with an optimized micro-kernel that keeps expert weights in standard layouts and instead fuses lightweight layout transforms into token gathering and stores; (2) applies NUMA-aware task partitioning under a simple page-interleaved policy without modifying the framework allocator; and (3) converts decode-phase memory matrix-vector operations into small matrix-matrix execution to utilize AMX. Across multiple MoE models, HiNa-MoE achieves up to 3.37x speedup for FFN kernels and up to 2.09x end-to-end inference speedup over state-of-the-art baselines, while remaining plug-and-play with existing frameworks and deployment workflows.
CommentsAccepted by PACT 2026