MonoMoE:一种用于量化MoE解码的高效融合超大内核
MonoMoE: An Efficient Fused Mega-kernel for Quantized MoE Decoding
浏览论文内容
中文总结 AI 辅助
MonoMoE是一种高效融合超大内核,通过优化块级量化MoE解码的运算流程,在NVIDIA H200 GPU上实现了MoE算子的显著加速,同时保持任务精度。
中文摘要 AI 辅助
混合专家(MoE)层在不按比例增加算术运算量的情况下提升模型容量,但它们的稀疏专家计算在自回归解码阶段难以高效执行。现有的分组和批量通用矩阵乘法(GEMM)采用令牌优先(token-major)模式:它们构建专家本地的令牌块,并从令牌维度获取并行性。当每个专家分配到的令牌数量较少时,这种组织方式会导致块填充和预处理,启动寿命较短的网格,从而未充分利用内存带宽,且将量化、激活和归约暴露为独立阶段。我们提出MonoMoE,一种用于块级量化MoE解码的权重优先持久超大内核。MonoMoE将完整解码步骤的令牌块置于细粒度张量核心N维度上,并在专家权重块上划分协作线程阵列(CTA),消除专家本地令牌物化并减少填充运算。持久网格将路由、top-k选择、量化、两个专家投影、激活和归约融合为一次启动;线程束专业化和就绪标志将辅助工作与主导专家权重流重叠。MonoMoE与vLLM集成,并通过生成的内核专业化和离线调度调优支持多种模型形状。在NVIDIA H200 GPU上,MonoMoE对完整路由MoE算子的加速比vLLM Triton分组GEMM最高达1.54倍,在不同模型和批量大小下比FlashMoE-FP8适配快2.20至3.84倍,在保持任务精度的同时,在评估的FP8模型上将每个输出令牌的端到端时间减少最多18.7%。MonoMoE的实现及配套构件开源,可在FlashInfer仓库获取。
英文摘要
Mixture-of-Experts (MoE) layers increase model capacity without proportionally increasing arithmetic, but their sparse expert computation is difficult to execute efficiently during autoregressive decode. Existing grouped and batched GEMMs are token-major: they construct expert-local token tiles and obtain parallelism from the token dimension. When few tokens reach each expert, this organization incurs tile padding and preprocessing, launches short-lived grids that underutilize memory bandwidth, and exposes quantization, activation, and reduction as separate stages. We present \textbf{MonoMoE}, a weight-major persistent megakernel for block-wise quantized MoE decode. MonoMoE places the complete decode-step token tile on the fine-grained tensor-core $N$ dimension and partitions CTAs over expert-weight tiles, eliminating expert-local token materialization and reducing padded arithmetic. A persistent grid fuses routing, top-$k$ selection, quantization, both expert projections, activation, and reduction in one launch; warp specialization and readiness flags overlap auxiliary work with the dominant expert-weight stream. MonoMoE is integrated with vLLM and supports multiple model shapes through generated kernel specializations and offline schedule tuning. On NVIDIA H200 GPUs, MonoMoE accelerates the complete routed-MoE operator by up to $\mathbf{1.54\times}$ over vLLM Triton Grouped GEMM, is $\mathbf{2.20}$--$\mathbf{3.84\times}$ faster than FlashMoE-FP8 adaptation across various models and batch sizes, and reduces end-to-end time per output token by up to $\mathbf{18.7\%}$ across the evaluated FP8 models, while preserving task accuracy. The MonoMoE implementation and supporting artifacts are open source and available in the \href{https://github.com/flashinfer-ai/flashinfer/tree/main/csrc/fused_moe/monomoe}{FlashInfer repository}.
发表机构
- Amazon AGI(亚马逊AGI)
机构由 AI 辅助整理,请以论文原文为准。