发表机构
University of Maryland, College Park(马里兰大学帕克分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对GPU平台上稀疏数据的集合通信,提出利用稀疏性算法及新格式Pici,能适应数据稀疏度并修改表示,在高输入稀疏度下相比NCCL有显著加速。
AI 中文摘要
高性能集合通信原语对各种高性能计算(HPC)和机器学习(ML)工作负载至关重要。现有如NCCL的库专为密集数据优化。本文针对全收集、归约散射和全归约三个集合通信引入利用稀疏性算法,实现新格式Pici,算法能适应稀疏度,在99%输入稀疏度下有显著加速。
英文摘要
Collective communication is essential to high performance computing and machine learning workloads, yet libraries such as NCCL do not exploit sparsity in message payloads. Sending only nonzero values can reduce network traffic, but explicitly handling sparsity introduces challenges such as compression and decompression overheads. We address these challenges with sparsity-exploiting versions of all-gather, reduce-scatter, and all-reduce collectives. Our implementations use a new bitvector-based format, Pici, designed for low space overhead and fast GPU-based compression and decompression. Further, our collective algorithms adapt to the degree of sparsity in data, modifying data representations during the course of the collective. At 99% input sparsity, our collectives achieve up to 5.25$\times$, 2.5$\times$, and 2.66$\times$ speedups over NCCL for all-gather, reduce-scatter, and all-reduce, respectively. Integrating our collectives into a representative deep learning application, we achieve a 26% end-to-end speedup.
CommentsAccepted at SC 2026