arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Weave:MoE 兆内核中用于计算-通信重叠的细粒度动态 SM 调度

Weave: Fine-Grained Dynamic SM Scheduling in an MoE Megakernel for Compute-Communication Overlap

Ziyu Huang, Yangjie Zhou, Chenhao Zhu, Peng Yu, Zihan Liu, Jinyu Liu, Shulai Zhang, Xingxun Tang, Hongzhe Yan, Xinhao Luo, Minyi Guo, Xiu Lin, Yinghao Yu, Guodong Yang, Liping Zhang, Shixuan Sun, Jingwen Leng

arXiv 2609.21483首次发表:更新:

发表机构

Shanghai Jiao Tong University; National University of Singapore; Fudan University; Alibaba Group(上海交通大学; 新加坡国立大学; 复旦大学; 阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对专家并行下 MoE 推理的通信开销,Weave 提出在持久兆内核中按层按 GPU 动态调度 SM,实现计算-通信重叠,显著提升性能。

AI 中文摘要

专家并行(EP)下的混合专家(MoE)推理将每个 MoE 层转变为分布式计算,并伴随昂贵的调度(dispatch)和合并(combine)通信。最先进的系统通过通信-计算重叠来降低此成本,将 GPU 的 SM 分别分配给通信和计算。然而,这种方法在空间和时间两个维度上仍会浪费 GPU 资源。在空间上,最佳的 SM 划分由每层的路由结果决定,并随层和 GPU 而变化,因此固定策略与工作负载不匹配,浪费 NVLink 带宽或计算吞吐量。在时间上,复杂的 MoE 数据依赖会引入气泡,导致 SM 空闲。我们提出 Weave,据我们所知,这是第一个执行细粒度动态 SM 调度的 MoE 重叠系统——在运行时根据路由结果逐层、逐 GPU 进行决策。一旦路由完成,每层的通信和计算量便已知;Weave 利用这种可预测性,通过运行在持久兆内核中的轻量级成本模型:空间调度器将 SM 划分为通信工作线程和计算工作线程,以匹配通信/计算吞吐量比;时间调度器协调两个工作线程组以最小化 SM 空闲。在 4x H100 SXM GPU 上,针对六个主流 MoE 模型,Weave 在 MoE 层上实现了 2.89 倍几何平均加速,在端到端上实现了 1.33 倍几何平均加速,优于五个最先进的基线。

英文摘要

Mixture-of-Experts (MoE) inference under expert parallelism (EP) turns each MoE layer into a distributed computation with costly dispatch and combine communication. State-of-the-art systems reduce this cost through communication-computation overlap, splitting the GPU's SMs for communication and computation respectively. However, this approach still leaves GPU resources wasted along two dimensions. Spatially, the best SM split is determined by each layer's routing result and varies across layers and GPUs, so fixed policies mismatch the workload and waste either NVLink bandwidth or compute throughput. Temporally, complex MoE data dependencies introduce bubbles that leave SMs idle. We present Weave, to our knowledge the first MoE overlap system that performs fine-grained dynamic SM scheduling - deciding per layer and per GPU by routing results at runtime. Once routing completes, each layer's communication and computation volumes become known; Weave exploits this predictability through a lightweight cost model running inside the persistent megakernel: a spatial scheduler partitions SMs into communication workers and computation workers to match the communication/computation throughput ratio, and a temporal scheduler coordinates the two worker groups to minimize SM idleness. On 4x H100 SXM GPUs across six mainstream MoE models, Weave achieves a 2.89x geometric-mean MoE-layer speedup and a 1.33x geometric-mean end-to-end speedup over five state-of-the-art baselines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑