arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GPU平台上用于动态和非结构化稀疏性的自适应空间高效集合通信

SpCCL: A Sparsity-Aware Collective Communication Library for GPU Platforms

Lannie Dalton Hough, Emir Gencer, Hoffmann Muki, Abhinav Bhatele

arXiv 2607.04676首次发表:更新:

发表机构

University of Maryland, College Park(马里兰大学帕克分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对GPU平台上稀疏数据的集合通信,提出利用稀疏性算法及新格式Pici,能适应数据稀疏度并修改表示,在高输入稀疏度下相比NCCL有显著加速。

AI 中文摘要

高性能集合通信原语对各种高性能计算(HPC)和机器学习(ML)工作负载至关重要。现有如NCCL的库专为密集数据优化。本文针对全收集、归约散射和全归约三个集合通信引入利用稀疏性算法,实现新格式Pici,算法能适应稀疏度,在99%输入稀疏度下有显著加速。

英文摘要

Collective communication is essential to high performance computing and machine learning workloads, yet libraries such as NCCL do not exploit sparsity in message payloads. Sending only nonzero values can reduce network traffic, but explicitly handling sparsity introduces challenges such as compression and decompression overheads. We address these challenges with sparsity-exploiting versions of all-gather, reduce-scatter, and all-reduce collectives. Our implementations use a new bitvector-based format, Pici, designed for low space overhead and fast GPU-based compression and decompression. Further, our collective algorithms adapt to the degree of sparsity in data, modifying data representations during the course of the collective. At 99% input sparsity, our collectives achieve up to 5.25$\times$, 2.5$\times$, and 2.66$\times$ speedups over NCCL for all-gather, reduce-scatter, and all-reduce, respectively. Integrating our collectives into a representative deep learning application, we achieve a 26% end-to-end speedup.

CommentsAccepted at SC 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑