arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

评估OpenMP Offloading在NVIDIA、AMD和Intel架构下的节点内多GPU编程应用:以3D热传导为例

Evaluating OpenMP Offloading for Intra-node Multi-GPU Programming across NVIDIA, AMD, and Intel Architectures: A 3D Heat Transfer Case Study

Ezhilmathi Krishnasamy

arXiv 2608.11882首次发表:更新:

发表机构

Rudolfovo-Science and Technology Centre Novo mesto; Faculty of Mechanical Engineering, University of Ljubljana; Department of Computer Science, FSTM, University of Luxembourg(鲁多尔沃夫诺沃梅斯托科学与技术中心; 卢布尔雅那大学机械工程学院; 卢森堡大学科学技术学院计算机科学系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究以3D热传导为例,对比分析OpenMP Offloading在三类GPU上的性能,发现其多GPU实现较单GPU可获2倍、4倍左右性能提升。

AI 中文摘要

当前,大多数超级计算机配备了NVIDIA、AMD或Intel等厂商的GPU,这些GPU能提供大量并行性和高吞吐量。单个计算节点(节点内)通常搭载多个GPU,一般为4个或更多。因此,有效利用单个计算节点内的所有GPU,对科学和工程领域的应用至关重要。然而,在将这些GPU用于科学计算前,需考虑多个因素,包括数据通信的实现、可用于跨这些GPU的编程模型,以及单个代码库在单个计算节点内不同GPU架构和配置下能达到的性能水平。OpenMP Offloading是一种基于指令的突出编程模型,可在NVIDIA、AMD和Intel这三类GPU上运行。本研究分析了使用OpenMP Offloading求解3D热方程的优势和性能挑战,3D热方程涉及主要计算、边界区域(halo)计算及通信。我们研究了其性能与原生GPU编程模型的差异,NVIDIA对应CUDA,AMD对应HIP,Intel对应SYCL。此外,我们证明,与单GPU的OpenMP Offloading实现相比,OpenMP Offloading在所有三类GPU上,使用2个GPU时性能提升约2倍,使用4个GPU时性能提升约4倍。该分析通过多种OpenMP Offloading实现系统地开展,这些实现采用不同的底层API进行内存分配、内存传输选项(同步、异步和对等),以及其他原生GPU编程模型,如NVIDIA的CUDA、AMD的HIP和Intel的SYCL。

英文摘要

Currently, most supercomputers are equipped with GPUs from manufacturers such as NVIDIA, AMD, or Intel, which provide substantial parallelism and high throughput. It is common for a single compute node (intranode) to host multiple GPUs, typically four or more. Therefore, effectively leveraging all these GPUs within a single compute node is essential for applications in scientific and engineering domains. However, several factors must be considered before utilizing these GPUs for scientific computing, including the implementation of data communication, the programming models available for use across these GPUs, and the level of performance that can be achieved with a single codebase across different GPU architectures and configurations within a single compute node. OpenMP Offloading is a prominent directive-based programming model that can be executed on all three GPU types: NVIDIA, AMD, and Intel. In this research, we present an analysis of the benefits and performance challenges of using OpenMP Offloading to address the 3D heat equations, which involve both primary computation, as well as halo computation and communication. For additional comparison and scalability study, we also consider the Conjugate Gradient method. We investigate how performance varies in relation to native GPU programming models-CUDA for NVIDIA, HIP for AMD, and SYCL for Intel. Furthermore, we demonstrate that OpenMP Offloading can achieve performance improvements of approximately 2x for 2 GPUs and around 4x for 4 GPUs when compared to single-GPU OpenMP Offloading implementations across all three GPU types. This analysis is conducted systematically through various OpenMP Offloading implementations that utilize different low-level APIs for memory allocation, memory transfer options (synchronous, asynchronous, and peer-to-peer), and other native GPU programming models such as CUDA(NVIDIA),HIP(AMD),and SYCL(Intel).

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑