mKernel:快速多GPU、多节点融合内核
mKernel: Fast Multi-GPU, Multi-Node Fused Kernels
- UC Berkeley(加州大学伯克利分校)
- UC Davis(加州大学戴维斯分校)
- UCLA(加州大学洛杉矶分校)
- University Politehnica of Bucharest(布加勒斯特理工大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
mKernel提出多GPU多节点融合内核库,通过SM划分、分层数据移动和GPU驱动网络,在块粒度重叠计算与通信,显著加速分布式训练与推理。
AI中文摘要:
通信已成为大规模模型分布式训练和推理的瓶颈。在单独流上以内核粒度将通信与计算重叠,仅能降低部分通信成本。融合内核通常通过在每个输出块生成后立即传输来获得更好的性能,但现有融合内核主要局限于单个NVLink域。我们提出mKernel,一个多GPU、多节点融合内核库,在块粒度上重叠计算、节点内NVLink通信和节点间RDMA。mKernel将持久内核的流式多处理器(SM)划分为计算和通信角色,并由片上GPU控制器在运行时自适应调整SM划分,因为最佳SM划分随内核和输入形状而变化。它分层组织数据移动,使穿越节点间网络的数据量最小化。最后,它通过轻量级命令队列和直接基于RDMA verbs实现的主机代理从GPU驱动网络,这使得相同内核能在任何网络后端(如InfiniBand和AWS EFA)上运行;我们惊讶地观察到,GPUDirect Async(IBGDA)相比主机辅助的GPU发起通信仅带来很少额外收益。我们实现了覆盖张量、序列和专家并行性的五个内核。在两个16-GPU H200集群上,mKernel在GEMM+AllReduce上实现高达1.72倍加速,在Ring Attention上实现1.88倍加速。
英文摘要:
Communication has become a bottleneck in distributed training and inference of large models. Overlapping communication with computation at the granularity of kernels, on separate streams, reduces only part of this communication cost. Fused kernels often have better performance by transmitting each output tile as soon as it is produced, but existing fused kernels are largely confined to a single NVLink domain. We present mKernel, a library of multi-GPU, multi-node fused kernels that overlap computation, intra-node NVLink communication, and inter-node RDMA at tile granularity. mKernel partitions the streaming multiprocessors (SMs) of a persistent kernel into compute and communication roles, and an on-GPU controller tunes the SM partition adaptively at run time, since the best SM partition varies with the kernel and the input shape. It structures data movement hierarchically so that data traversing the inter-node network is minimized. Finally, it drives the network from the GPU through a lightweight command queue and host proxy implemented directly on RDMA verbs, which allows the same kernels to run on any network backend (e.g. InfiniBand and on AWS EFA); we observe, surprisingly, that GPUDirect Async (IBGDA) yields little additional benefit over host-assisted GPU-initiated communication. We implement five kernels spanning tensor, sequence, and expert parallelism. On two 16-GPU H200 clusters, mKernel achieves speedups of up to 1.72x on GEMM+AllReduce and $1.88\times$ on Ring Attention.