arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

每微秒都至关重要:在GPU集合通信中实现接近光速的延迟

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives

Siyuan Shen, Anton Korzh, John Bachan, Tiancheng Chen, Arnav Goel, Ludwig Schneider, Pouya Kousha, Zhenhao He, Sylvain Jeaugey, Kamil Iskra, Nishank Chandawala, Jeff R. Hammond, Torsten Hoefler

arXiv 2607.16100首次发表:更新:

AI 中文总结

研究在扩展网络中使GPU集合通信接近硬件光速下限的方法,确定关键原则,基于NCCL开发低延迟接口实现新对称集合,微基准测试和实际应用表明能降低延迟、提升吞吐量,对AI推理和传统HPC工作负载有益。

AI 中文摘要

GPU集合通信通常针对带宽进行优化,但许多新兴工作负载越来越受延迟限制。长上下文解码量大的大语言模型推理就是典型例子,服务大型模型需多个GPU,许多小集合直接处于令牌生成的关键路径上。因此,即使微秒级的开销也会影响性能和成本。本文研究如何在扩展网络中接近GPU集合通信的硬件光速下限。确定了接近最优设计的关键原则,基于NCCL的设备端API开发了低延迟接口来构建自定义集合内核,并在NCCL中实现了新的对称集合。微基准测试表明中小消息的延迟大幅降低,开销降至绝对光速下限的7%以内。集成到实际应用中,这些内核改善了大语言模型推理中的令牌间延迟和吞吐量,并加速了cuSOLVERMp,对人工智能推理和传统HPC工作负载都有好处。

英文摘要

GPU collective communication is typically optimized for bandwidth, yet many emerging workloads are increasingly limited by latency. Long-context decode-heavy large language model (LLM) inference is a prime example, where serving large models requires multiple GPUs, and many small collectives lie directly on the critical path of token generation. Therefore, even microsecond of overhead can impact performance and cost. In this work, we study how to approach the hardware Speed-of-Light (SoL) lower bound for GPU collectives within a scale-up network. We identify key principles for near-optimal designs, including barrier-free synchronization and efficient use of symmetric memory and multicast. Building on NCCL's device-side API, we develop low-latency interfaces for constructing custom collective kernels and use them to implement new symmetric collectives in NCCL. Microbenchmarks show substantial latency reductions for small and medium messages, reducing overhead to within 7% of the absolute SoL lower bound. When integrated into real applications, these kernels improve inter-token latency and throughput in LLM inference and accelerate cuSOLVERMp, demonstrating benefits for both AI inference and traditional HPC workloads.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑