发表机构
Linnaeus University(林纳大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对专家混合模型在多GPU系统中通信延迟问题,提出通过瓦片级信号与调度实现细粒度计算-通信重叠的方法,结合生产者-消费者协同设计,在4个A100 GPU平台评估中取得加速效果,提升了不同场景下的性能。
AI 中文摘要
专家混合(MoE)架构在不按比例增加计算成本的情况下提高了模型容量,成为将大语言模型扩展到万亿参数规模的关键构建模块。其有效部署依赖于跨多个GPU的分布式执行,每个MoE层涉及两次全对全通信。传统实现方法在专家计算完成后才启动返回全对全通信,导致关键路径上的通信延迟并降低GPU利用率。本文提出一种细粒度方法,通过瓦片级信号与调度使专家计算与第二次全对全通信重叠。我们的生产者-消费者协同设计结合了:(1)一个持久的按秩计算内核(生产者),覆盖该秩上的所有本地专家,以消除重复内核启动开销并优先处理远程关键瓦片;(2)一个在流多处理器(SM)的小专用分区上的持久通信内核(消费者),在瓦片准备好时发出段粒度传输。我们的协同设计避免了对底层计算算子或通信原语的侵入性更改,使其在提高多GPU系统上分布式MoE执行效率方面具有实用性。在4个A100 GPU平台上,针对3个MoE模型与4个现有最佳MoE系统进行评估,我们的方法实现了高达2.64倍的端到端加速和2.74倍的MoE层加速。与传统的非重叠基线相比,我们的方法在不同的通用矩阵乘法(GEMM)形状、路由器模式以及广泛的生产者/消费者SM分区中,始终提高了算子级和MoE层级的性能,同时保持正确性。
英文摘要
Mixture-of-Experts (MoE) architectures increase model capacity without proportionally increasing computation cost and have become a key building block for scaling large language models (LLMs) to trillion-parameter regimes. Efficient deployment of these MoE models relies on distributed execution across multiple GPUs, where each MoE layer involves two all-to-all communications: dispatching tokens to expert ranks and returning the expert outputs to their source ranks. Conventional MoE implementations launch this return all-to-all after expert compute completes, exposing communication latency on the critical path and reducing GPU utilization. We present a fine-grained approach that overlaps expert compute with the second all-to-all via tile-level signaling and scheduling. Our producer-consumer co-design combines: (1) a persistent per-rank computation kernel (producer) that covers all local experts on the rank to eliminate repeated kernel launch overhead and prioritizes remote-critical tiles, and (2) a persistent communication kernel (consumer) on a small dedicated partition of streaming multiprocessors (SMs) that issues segment-granular transfers as tiles become ready. Our co-design avoids intrusive changes to the underlying computation operators or communication primitives, making it practical for improving distributed MoE execution efficiency on multi-GPU systems. On a 4-A100 GPU platform, evaluated on three MoE models against four state-of-the-art MoE systems, our approach achieves up to 2.64x end-to-end speedup and 2.74x MoE-layer speedup. Compared with a conventional non-overlap baseline, our approach consistently improves both operator- and MoE-layer-level performance across varying GEMM shapes, router modes, and a broad range of producer/consumer SM partitions, while preserving correctness.
CommentsTo appear at the 55th International Conference on Parallel Processing (ICPP 26)