arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.07536cs.DC

面向细粒度MoE计算-通信重叠的解析资源管理

Analytical Resource Management for Fine-grained MoE Computation-Communication Overlap

Hongyu Liu, Minyu Cui, Miquel Pericas

首次发表
浏览论文内容

中文总结 AI 辅助

针对分布式MoE推理中的细粒度计算-通信重叠,提出波量化解析模型和启动时资源管理器,动态选择通信CTA数量和资源分区,在A100上实现最高4.2倍加速。

中文摘要 AI 辅助

分布式混合专家(MoE)推理中的细粒度计算-通信重叠允许在部分计算结果就绪时即开始通信。然而,执行计算和通信的协作线程数组(CTA)会竞争流式多处理器(SM)上有限的驻留容量。由于驻留的CTA通常保留其分配的SM资源直至完成,无法共驻的CTA必须等待资源,导致波状执行。固定的资源分区无法适应输入大小、路由专家负载和内核配置的变化,可能导致通信积压或降低专家计算并行度。我们提出了一种波量化解析模型和启动时资源管理器,用于依赖耦合的重叠流水线。利用路由块计数、内核占用率、GPU驻留约束和分裂级就绪依赖,它在每次启动前选择通信CTA数量和资源分区,无需候选执行、逐工作负载剖析或内核重编译。我们将该方法集成到FLUX中公开的COMET A100实现中。我们在四块NVIDIA A100 GPU上,在多种并行策略下,评估了三个MoE模型,涵盖GEMM2+GatherRS算子、完整的后路由MoE层和完整模型预填充级别。在15个真实p90工作负载中,解析选择器相对于测量得到的oracle实现了3.22%的平均遗憾,平均求解器开销为0.157微秒。与COMET相比,我们的方法在GEMM2+GatherRS算子、完整的后路由MoE层和完整模型预填充上分别实现了2.528倍、1.771倍和1.185倍的几何平均加速,最大加速比分别为4.218倍、2.584倍和1.439倍。在每一个可行的TP=2/EP=2序列长度(至少4096)下,我们的实现均优于COMET、Megatron core-TE和FastMoE TP+NCCL。

英文摘要

Fine-grained computation--communication overlap in distributed Mixture-of-Experts (MoE) inference allows communication to begin as partial compute results become ready. However, cooperative thread arrays (CTAs) performing computation and communication contend for finite residency capacity on streaming multiprocessors (SMs). Because a resident CTA generally retains its allocated SM resources until completion, CTAs that cannot be co-resident must wait for resources, resulting in wave-like execution. A fixed resource partition cannot adapt to changes in input size, routed expert load, and kernel configuration, potentially causing a communication backlog or reducing expert compute parallelism. We present a wave-quantized analytical model and launch-time resource manager for dependency-coupled overlap pipelines. Using routed-tile counts, kernel occupancy, GPU residency constraints, and split-level readiness dependencies, it selects the communication-CTA count and resource partition before each launch without candidate execution, per-workload profiling, or kernel recompilation. We integrate the method into the public COMET A100 implementation in FLUX. We evaluate three MoE models on four NVIDIA A100 GPUs under several parallelism strategies at the GEMM2+GatherRS operator, complete post-router MoE layer, and complete-model prefill levels. Across 15 real-p90 workloads, the analytical selector achieves 3.22 percent mean regret relative to the measured oracle with a mean solver overhead of 0.157 microseconds. Over COMET, our method achieves geometric-mean speedups of 2.528x at the GEMM2+GatherRS operator, 1.771x at the complete post-router MoE layer, and 1.185x for complete-model prefill, with maxima of 4.218x, 2.584x, and 1.439x, respectively. At every feasible TP=2/EP=2 sequence length of at least 4,096, our implementation outperforms COMET, Megatron core-TE, and FastMoE TP+NCCL.

发表机构

  • Chalmers University of Technology(查尔姆斯理工大学)
  • University of Gothenburg(哥德堡大学)
  • Linnaeus University(林奈大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑