每个专家都重要:用于内存高效W4A16推理的ExactMoE
Every Expert Counts: ExactMoE for Memory-Efficient W4A16 Inference
浏览论文内容
中文总结 AI 辅助
ExactMoE仅量化路由专家并结合缓存与融合内核,大幅降低MoE模型推理的GPU内存占用,同时保留高吞吐量与准确率,达成内存效率与性能的良好平衡。
中文摘要 AI 辅助
稀疏混合专家(MoE)语言模型通过仅为每个token激活一小部分专家来减少运算量,但部署时仍需存储和移动完整的专家库。我们提出ExactMoE,这是一种推理设计,仅对路由后的专家应用对称的分组128四比特权重量化,将这些专家以内核原生的MARLIN形式存储在固定主机内存中,并通过可配置的GPU驻留插槽缓存和融合分组MoE内核执行所有选定的专家。路由层、注意力机制、嵌入层、归一化层和语言模型头仍保持为BF16精度。“Exact”指专家完全可用且top-k路由过程不变:没有专家被修剪、替换或强制在CPU上执行,并不意味着与BF16模型的数值完全一致。在OLMoE-1B-7B-0924-Instruct上,于单个NVIDIA L4评估,16插槽配置将峰值预留GPU内存从14.168 GiB降至1.836 GiB(减少87.04%),同时保留BF16解码吞吐量的81.85%;完全驻留的64插槽配置达到31.923 tokens/s,而BF16为21.662 tokens/s,同时预留4.061 GiB。在12450个零样本多选问题上,ExactMoE获得70.3534%的归一化准确率,而BF16为70.8996%,保留了基线准确率的99.23%。在匹配的16-token消融实验中,融合分组执行的速度是顺序W4参考的1.97倍。这些结果确定了完整专家MoE推理的实用内存-传输-吞吐量前沿。
英文摘要
Sparse mixture-of-experts (MoE) language models reduce arithmetic by activating only a small subset of experts per token, yet deployment still requires storing and moving the full expert bank. We present ExactMoE, an inference design that applies symmetric group-128 four-bit weight quantization only to routed experts, stores those experts in kernel-native MARLIN form in pinned host memory, and executes all selected experts through a configurable GPU-resident slot cache and fused grouped MoE kernels. The router, attention, embeddings, normalization layers, and language-model head remain in BF16. "Exact" refers to complete expert availability and an unchanged top-k routing procedure: no expert is pruned, substituted, or forced to execute on the CPU. It does not imply numerical identity with the BF16 model. On OLMoE-1B-7B-0924-Instruct, evaluated on a single NVIDIA L4, a 16-slot configuration reduces peak reserved GPU memory from 14.168 to 1.836 GiB (87.04%) while retaining 81.85% of BF16 decode throughput. A fully resident 64-slot configuration reaches 31.923 tokens/s versus 21.662 tokens/s for BF16 while reserving 4.061 GiB. Across 12,450 zero-shot multiple-choice questions, ExactMoE obtains 70.3534% normalized accuracy versus 70.8996% for BF16, retaining 99.23% of the baseline accuracy. In a matched 16-token ablation, fused grouped execution is 1.97x as fast as a sequential W4 reference. These results identify a practical memory-transfer-throughput frontier for complete-expert MoE inference.
发表机构
- Appendture(阿彭彻(Appendture))
机构由 AI 辅助整理,请以论文原文为准。