arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Entwine:协调GPU间的分块计算与细粒度通信

Entwine: Coordinating Tiled Computation and Fine-Grained Communication across GPUs

Kai Ma, Quanfeng Lv, Jingguo Ge, Bowei Dai, Kefan Ruan

arXiv 2609.11562首次发表:更新:

发表机构

State Key Laboratory of Cyberspace Security Defense; Institute of Information Engineering, Chinese Academy of Sciences; University of Chinese Academy of Sciences; Institute of Microelectronics, Chinese Academy of Sciences(网络空间安全防御国家重点实验室; 中国科学院信息工程研究所; 中国科学院大学; 中国科学院微电子研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Entwine通过协调分块计算顺序、细粒度通信和SM资源分配,实现GPU间计算与通信的高效重叠,在张量并行LLM负载上相比cuBLAS+NCCL取得1.232倍几何平均加速。

AI 中文摘要

现代高性能GPU计算将张量划分为分块(tile)以利用数据重用和并行性。单个分块的计算比整个张量的计算完成得更早,从而创造了重叠计算与通信的机会。然而,计算与通信进度之间的不匹配会限制这些机会。当没有数据准备好时,通信会停滞,而当数据突发到达时,通信可能会滞后。通信还可能通过消耗共享资源来减慢计算速度,从而抵消重叠带来的收益。我们提出了Entwine,它协调分块计算顺序、细粒度通信和SM资源分配,以最小化整体完成时间。Entwine重新排序分块计算,以更规律的节奏为通信生成数据。Entwine将此调度与基于SM的细粒度通信相结合,以低延迟和低开销处理分块结果。由于通信内核也消耗SM资源,Entwine协调它们的分配,以在通信进度与计算减速之间取得平衡。在具有代表性的张量并行LLM工作负载中,Entwine相比cuBLAS+NCCL实现了1.232倍(最高1.433倍)的几何平均加速,并在几何平均上比最先进的重叠基线高出3.1%-9.8%。我们将在发表后开源我们的实现。

英文摘要

Modern high-performance GPU computations partition tensors into tiles to exploit data reuse and parallelism. Individual tile computations complete earlier than the full tensor computation, creating opportunities to overlap computation and communication. However, a mismatch between computation and communication progress can limit these opportunities. Communication stalls when no data is ready, and may lag when data arrives in bursts. Communication can also slow computation by consuming shared resources, offsetting the benefits of overlap. We present Entwine, which coordinates tile computation order, fine-grained communication, and SM resource allocation to minimize overall completion time. Entwine reorders tile computation to produce data for communication at a more regular pace. Entwine couples this schedule with fine-grained SM-based communication to process tile results with low latency and low overhead. Since the communication kernel also consumes SM resources, Entwine coordinates their allocation to balance communication progress against computation slowdown. Across representative tensor-parallel LLM workloads, Entwine achieves a geomean speedup of 1.232x (up to 1.433x) over cuBLAS+NCCL, and outperforms state-of-the-art overlap baselines by 3.1-9.8% in geomean. We will open-source our implementation upon publication.

Comments15 pages, including references and appendices

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑