arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

T-CCL:利用张量内存加速器实现资源高效且高性能的集合通信

T-CCL: Resource Efficient and Performant Collective Communication using Tensor Memory Accelerator

Keyvan Dadashzadeh, Yuehong Zhou, Minyu Cui, Miquel Pericas

arXiv 2610.07098首次发表:更新:

发表机构

Chalmers University of Technology; University of Gothenburg; Linnaeus University(查尔姆斯理工大学; 哥德堡大学; 林奈大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

T-CCL利用张量内存加速器卸载数据移动和归约,以流水线异步操作减少SM占用,实现资源高效集合通信,相比NCCL最高提速3.42倍,并提升重叠计算与推理吞吐。

AI 中文摘要

基于大型Transformer的模型越来越依赖于多GPU执行,这需要GPU之间频繁的集合通信。现有的通信库通常依赖大量GPU线程来实现高带宽或低延迟,导致流式多处理器(SM)侧资源占用较大。这种占用可能限制其他GPU工作可用的资源,尤其是在通信和计算并发执行时。因此,高效的集合通信不仅应实现高集合性能,还应减少其SM侧资源使用。本文提出了T-CCL,一种基于张量内存加速器(TMA)的资源高效集合通信库,用于节点内通信。T-CCL将数据移动和归约操作卸载到TMA,并将每个集合操作执行为一系列流水线化的异步TMA操作,从而在保持高带宽的同时减少集合通信所需的SM资源。在AllReduce、AllGather和ReduceScatter集合操作上的评估中,T-CCL在通信资源不受限制时比NCCL性能提升高达2.4倍,在资源受限预算下提升高达3.42倍,与NCCL最新的对称内存内核保持竞争力,并且在分析案例中占用相同或更少的SM。在GEMM-集合重叠案例研究中,将通信后端从NCCL切换到T-CCL,使得两个GPU上相对于顺序基线的平均算子级加速从1.12倍提高到1.25倍,四个GPU上从1.04倍提高到1.14倍,因为T-CCL使用更少的SM进行通信,为重叠的GEMM留下更多可用SM。集成到vLLM作为通信后端后,T-CCL相比vLLM的自动后端调度,端到端推理吞吐量提升高达1.31倍,在对话和重解码工作负载的每个评估批次大小上均优于后者。

英文摘要

Large transformer-based models increasingly depend on multi-GPU execution, which requires frequent collective communication among GPUs. Existing communication libraries often rely on many GPU threads to achieve high bandwidth or low latency, resulting in a large streaming multiprocessor (SM)-side resource footprint. This footprint can limit the resources available to other GPU work, particularly when communication and computation execute concurrently. Thus, efficient collective communication should not only achieve high collective performance but also reduce its SM-side resource usage. This paper presents T-CCL, a resource-efficient collective communication library based on the Tensor Memory Accelerator (TMA) for intra-node communication. T-CCL offloads both data movement and reduction operations to TMA and executes each collective as a pipelined series of asynchronous TMA operations, reducing the SM resources required for collective communication while maintaining high bandwidth. Evaluated across AllReduce, AllGather, and ReduceScatter collectives, T-CCL outperforms NCCL by up to 2.4x with unrestricted communication resources and up to 3.42x under restricted resource budgets, remains competitive with NCCL's recent symmetric-memory kernels, and occupies the same or fewer SMs in profiled cases. In a GEMM-collective overlap case study, switching the communication backend from NCCL to T-CCL raises the average operator-level speedup over a sequential baseline from 1.12x to 1.25x on two GPUs and from 1.04x to 1.14x on four GPUs, as T-CCL uses fewer SMs for communication, leaving more SMs available to the overlapped GEMM. Integrated into vLLM as a communication backend, T-CCL improves end-to-end inference throughput over vLLM's automatic backend dispatch by up to 1.31x, outperforming it at every evaluated batch size on both the conversation and decode-heavy workloads.

CommentsWorkshops on Supercomputing (SC'26)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑