AI 中文总结
本文设计实现了基于NVIDIA NVSHMEM的多GPU并发队列系统,以Bellman-Ford SSSP算法为案例测试,在A100 GPU及10种测试图上分别取得最高3.92倍、3.03倍的加速比,为同类首个多GPU并发队列实现。
AI 中文摘要
本文提出并实现了一种基于NVIDIA NVSHMEM的多GPU并发队列系统,采用Bellman-Ford算法作为案例研究来评估所提出的并发FIFO队列的性能,该多GPU实现是已知的同类首个实例。实验结果表明,在四块NVIDIA A100 GPU上,该多GPU队列实现相较于单GPU基准实现,最高加速比达3.92倍,平均加速比为3.04倍;将其应用于Bellman-Ford单源最短路径(SSSP)算法时,在取自SuiteSparse矩阵集合的10种不同类型图上测试,该多GPU系统相较于单GPU实现,最高加速比达3.03倍,平均加速比为2.65倍。
英文摘要
This paper presents the design and implementation of a multi-GPU concurrent queue system using NVIDIA's NVSHMEM. The Bellman-Ford algorithm is used as a case study to evaluate the performance of the proposed concurrent FIFO queue, with this multi-GPU implementation being the first known instance of its kind. Experimental results demonstrate that the multi-GPU queue implementation achieves a maximum speedup of 3.92x and an average speedup of 3.04x over the singleGPU baseline on four NVIDIA A100 GPUs. When applied to the Bellman-Ford Single-Source Shortest Path (SSSP) algorithm, the multi-GPU system achieves a maximum speedup of 3.03x and an average speedup of 2.65x compared to the single-GPU implementation, tested on 10 graphs of different kinds taken from the SuiteSparse Matrix Collection.
CommentsIEEE ISPDC 2026 published conference paper
DOI:10.1109/ISPDC69862.2026.00016