AI 中文总结
本文针对NVIDIA、AMD、Intel三大GPU厂商,探究Fortran 'do concurrent'循环在GPU加速Fortran应用中的可移植性,发现纯Fortran代码可被GPU加速,补充指令可提升性能,相关技术正快速进步。
AI 中文摘要
人们对使用标准语言结构开展并行与加速的高性能计算(HPC)的兴趣持续增长,以避免依赖(有时是特定厂商的)外部API。对于Fortran应用而言,'do concurrent'循环这类语言特性,让编译器仅通过标准语言就能实现多线程、GPU加速甚至分布式多节点代码成为可能。本文针对三大GPU厂商(NVIDIA、AMD和Intel),探究将'do concurrent'用于GPU加速Fortran应用的当前状态。我们采用一款生产级应用测试其当前能力,明确仅使用标准语言可实现的场景,以及仍需或必须补充基于指令的API(如OpenMP)的场景;借助GPU感知MPI库开展多GPU测试。研究发现,三大GPU厂商如今均可对纯Fortran(零指令)代码进行GPU加速,但手动数据移动指令可助力提升性能与兼容性。结果表明,借助Fortran标准语言实现GPU加速科学HPC代码的性能可移植性,相关技术正快速进步。
英文摘要
There continues to be growing interest in using standard language constructs for parallel and accelerated HPC computing, avoiding the need for (sometimes vendor-specific) external APIs. For Fortran applications, language features such as 'do concurrent' loops open the door for compilers to implement multi-threaded, GPU-accelerated, and even distributed multi-node code with only the standard language. Here, we explore the current status of using 'do concurrent' for GPU-accelerated Fortran applications across three major GPU vendors (NVIDIA, AMD, and Intel). Using a production application, we test their current capabilities, showing where the standard language alone can be used, and where augmenting the code with a directive-based API (e.g., OpenMP) is still desirable or required. Multi-GPU tests are performed with GPU-aware MPI libraries. We find that the three GPU vendors can now GPU-accelerate pure Fortran (zero directives), but that manual data movement directives can help with performance and compatibility. The results show that there is rapid advancement towards making GPU-accelerated scientific HPC code performance portable using the Fortran standard language.
Comments12 pages, 6 figures