arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向多GPU平台的并发队列系统:应用于Bellman-Ford单源最短路径算法

A Concurrent Queue System for Multi-GPU Platforms: Application to Bellman-Ford SSSP

Beyza Cavusoglu

arXiv 2608.21826首次发表:更新:

AI 中文总结

本文设计实现了基于NVIDIA NVSHMEM的多GPU并发队列系统,以Bellman-Ford SSSP算法为案例测试,在A100 GPU及10种测试图上分别取得最高3.92倍、3.03倍的加速比,为同类首个多GPU并发队列实现。

AI 中文摘要

本文提出并实现了一种基于NVIDIA NVSHMEM的多GPU并发队列系统,采用Bellman-Ford算法作为案例研究来评估所提出的并发FIFO队列的性能,该多GPU实现是已知的同类首个实例。实验结果表明,在四块NVIDIA A100 GPU上,该多GPU队列实现相较于单GPU基准实现,最高加速比达3.92倍,平均加速比为3.04倍;将其应用于Bellman-Ford单源最短路径(SSSP)算法时,在取自SuiteSparse矩阵集合的10种不同类型图上测试,该多GPU系统相较于单GPU实现,最高加速比达3.03倍,平均加速比为2.65倍。

英文摘要

This paper presents the design and implementation of a multi-GPU concurrent queue system using NVIDIA's NVSHMEM. The Bellman-Ford algorithm is used as a case study to evaluate the performance of the proposed concurrent FIFO queue, with this multi-GPU implementation being the first known instance of its kind. Experimental results demonstrate that the multi-GPU queue implementation achieves a maximum speedup of 3.92x and an average speedup of 3.04x over the singleGPU baseline on four NVIDIA A100 GPUs. When applied to the Bellman-Ford Single-Source Shortest Path (SSSP) algorithm, the multi-GPU system achieves a maximum speedup of 3.03x and an average speedup of 2.65x compared to the single-GPU implementation, tested on 10 graphs of different kinds taken from the SuiteSparse Matrix Collection.

CommentsIEEE ISPDC 2026 published conference paper

DOI:10.1109/ISPDC69862.2026.00016

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑