AI 中文总结
研究MoE推理,提出分布感知框架,结合有效专家度量与反向建模过程生成可控分布。还介绍DA - MoE运行时,能匹配实时路由直方图与离线分布选最优内核,在HumanEval - X服务跟踪中提升了融合MoE延迟的几何平均值。
AI 中文摘要
混合专家(MoE)推理由形状随运行时路由分布变化的稀疏专家通用矩阵乘法(GEMM)组成。现有服务系统通常使用静态令牌计数桶来选择融合MoE内核,忽略了决定切片填充、内存复用和内核效率的每个专家的路由分布。我们引入了一个用于建模和基准测试MoE推理的分布感知框架。该框架将紧凑的有效专家度量与基于狄利克雷的反向建模过程相结合,以生成用于系统硬件研究的可控路由分布。通过它,我们表明最佳融合MoE内核会随路由偏差和令牌计数而变化。我们还提出了DA - MoE,这是一种用于NVIDIA GPU的驻留在GPU上的内核调度运行时,它将实时路由直方图与离线调整的分布相匹配,并在没有CPU - GPU同步的情况下选择接近最优的融合MoE内核。在HumanEval - X服务跟踪中,DA - MoE在DeepSeek - V3上可将融合MoE延迟的几何平均值提高1.16倍,在Kimi K2上提高1.29倍,峰值加速比分别为1.40倍和1.56倍。
英文摘要
Mixture-of-Experts (MoE) inference consists of sparse expert GEMMs whose shapes vary with the runtime routing distribution. Existing serving systems typically select fused-MoE kernels using static token-count buckets, ignoring the per-expert routing distribution that determines tile padding, memory reuse, and kernel efficiency. We introduce a distribution-aware framework for modeling and benchmarking MoE inference. The framework combines the compact Effective Experts metric with a Dirichlet-based reverse-modeling procedure that generates controllable routing distributions for systematic hardware studies. Using it, we show that the best fused-MoE kernel changes with routing skew and token count. We further present DA-MoE, a GPU-resident kernel-dispatch runtime for NVIDIA GPUs that matches the live routing histogram to offline-tuned distributions and selects a near-optimal fused-MoE kernel without CPU--GPU synchronization. On HumanEval-X serving traces, DA-MoE improves geomean fused-MoE latency by 1.16X on DeepSeek-V3 and 1.29X on Kimi K2, with peak speedups of 1.40X and 1.56X.
Comments13 pages, 15 figures