arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GPU发起的通信:深入剖析至核心

GPU-Initiated Communication: Dissecting Down to the Bone

Javid Baydamirli, Ismayil Ismayilov, Kaan Oktay, Didem Unat

arXiv 2610.01380首次发表:更新:

发表机构

Koç University; fal(科奇大学; fal)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文深入剖析GPU发起的通信机制,通过最小化实现和多种库对比,揭示提交路径、队列共享及资源成本对性能的影响,指出仅靠提交路径无法预测通信性能。

AI 中文摘要

GPU发起的通信允许GPU线程直接向网卡(NIC)提交远程直接内存访问(RDMA)操作。该机制支撑着NVSHMEM、NCCL GIN和DeepEP,这些库服务于混合专家(MoE)模型中细粒度、延迟敏感的通信,然而其性能特性和优化在源代码之外鲜有文档记录,且库之间的比较未能将硬件机制的成本与围绕它的库的成本分离开来。本文在GPU-NIC边界对GPU发起的通信进行深入剖析。我们首先详细阐述GPU侧的网络路径:队列放置、工作请求构造、门铃排序和完成语义。随后,我们介绍mini-gda和mini-proxy,这是针对GPU和CPU代理提交路径的最小化传输实现,并在NVIDIA H100、H200、B200和GB200平台上与NVSHMEM IBGDA、NCCL GIN、DeepEP、UCCL-EP、MSCCL++和fabric-lib一同进行测量。一条最小化的GPU路径发起一个操作需0.7微秒,完成需4.0微秒;库通过队列管理、内存排序和完成范围增加了高达4.6微秒的发起时间,且发起时间随SM时钟频率变化。在空闲状态下,调优后的CPU代理可匹配或超越GPU路径的性能,但代价是需要一个专用核心,其运行状态决定了延迟和吞吐量。在任一路径上,与批量流量共享队列会将延迟提高一到三个数量级。要达到我们InfiniBand平台260 M msg/s的上限,需要门铃批处理和队列并行,而这两者都有资源成本:通信代码即使未被使用也可能减少GPU块驻留,且全对全流量在约3000个活动连接时损失其NIC消息速率的59%。因此,仅提交路径并不能预测通信性能。我们的实验代码和结果可在以下网址获取:https://this https URL。

英文摘要

GPU-initiated communication lets GPU threads post RDMA operations directly to the NIC. It underpins NVSHMEM, NCCL GIN, and DeepEP, which serve the fine-grained, latency-critical communication of Mixture-of-Experts (MoE) models, yet its performance characteristics and optimizations remain scarcely documented beyond source code, and library comparisons fail to separate the costs of the hardware mechanism from those of the library around it. This paper dissects GPU-initiated communication at the GPU-NIC boundary. We first detail the GPU-side network path: queue placement, work-request construction, doorbell ordering, and completion semantics. We then introduce mini-gda and mini-proxy, minimal transports for the GPU and CPU-proxy submission paths, and measure them alongside NVSHMEM IBGDA, NCCL GIN, DeepEP, UCCL-EP, MSCCL++, and fabric-lib on NVIDIA H100, H200, B200, and GB200 platforms. A minimal GPU path issues an operation in 0.7 $μ$s and completes in 4.0 $μ$s; libraries add up to 4.6 $μ$s of issue time through queue management, memory ordering, and completion scope, and issue time scales with the SM clock. A tuned CPU proxy matches or beats the GPU path at idle, at the cost of a dedicated core whose operating state sets its latency and throughput. On either path, sharing a queue with bulk traffic raises latency by one to three orders of magnitude. Reaching the 260 M msg/s ceiling of our InfiniBand platform requires doorbell batching and queue parallelism, and both have resource costs: communication code can reduce GPU block residency even when unused, and all-to-all traffic loses 59% of its NIC message rate at about 3,000 active connections. The submission path alone therefore does not predict communication performance. Our experiment code and results are available at https://github.com/ParCoreLab/Dissecting-GPU-Communication-Experiments.

Comments15 pages, 11 figures, 16 tables. Code and data: https://github.com/ParCoreLab/Dissecting-GPU-Communication-Experiments

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑