迈向基于单一 Fortran 代码库的异构百亿亿次 CFD:高阶非结构化求解器 UCNS3D 的 OpenMP 目标加速
Towards Heterogeneous Exascale CFD with a Single Fortran Code Base: OpenMP Target Acceleration of the High-Order Unstructured Solver UCNS3D
- Cranfield University(克兰菲尔德大学)
- The University of Edinburgh(爱丁堡大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文通过单一Fortran代码库的OpenMP目标现代化改造,实现了高阶非结构化CFD求解器UCNS3D在异构百亿亿次系统上的加速,验证了数值精度并显著提升了性能。
AI中文摘要:
异构百亿亿次系统正在重塑计算流体力学,然而为加速器重写成熟的高阶求解器可能会导致代码库碎片化并削弱可重复性。我们提出了对 UCNS3D(一种用于可压缩流的高阶非结构化有限体积求解器)的单一源代码 OpenMP 目标现代化改造。通用的 CPU/GPU 实现保留了其数值公式、Fortran 代码库和工作流程。我们并非移植孤立的核函数,而是启用了完整的显式时间推进路径:高阶重构、梯度评估、无粘和粘性通量、边界处理、晕环交换以及解更新。该实现采用了持久目标数据区域、扁平运行时数组、编译时临时存储大小设定、显式局部代数、可目标调用的例程以及具有设备驻留缓冲区的 GPU 感知 MPI。这些选择解决了高阶非结构化 CFD 的不规则模板、复杂数据结构以及大量临时存储问题。使用可压缩 Taylor-Green 涡的验证表明,CPU 和 GPU 的耗散历史在 LUMI 上以及已发表的参考数据均达到机器精度。在 ARCHER2 GPU 平台上的单节点运行提供了独立的可移植性检查。端到端性能在完全填充的节点上使用 Taylor-Green 涡、LM1021 音爆和 NASA 高升力 CRM 基准进行了评估。相对于之前的生产实现,重构后的 CPU 路径速度提高了 1.27-1.67 倍,而 GPU 卸载实现了同节点 2.71-4.05 倍的加速以及 84-101% 的强扩展效率。结果表明,基于标准的 OpenMP 提供了可移植的、生产规模的加速,而不会牺牲数值保真度、CPU 性能或软件可持续性。
英文摘要:
Heterogeneous exascale systems are reshaping computational fluid dynamics, yet rewriting mature high-order solvers for accelerators can fragment code bases and weaken reproducibility. We present a single-source OpenMP target modernisation of UCNS3D, a high-order unstructured finite-volume solver for compressible flows. The common CPU/GPU implementation preserves its numerical formulation, Fortran code base, and workflows. Rather than porting isolated kernels, we enable the complete explicit time-advancement path: high-order reconstruction, gradient evaluation, inviscid and viscous fluxes, boundary treatment, halo exchange, and solution update. The implementation uses persistent target data regions, flat run-time arrays, compile-time sizing of temporary storage, explicit local algebra, target-callable routines, and GPU-aware MPI with device-resident buffers. These choices address the irregular stencils, complex data structures, and substantial temporary storage of high-order unstructured CFD. Verification using the compressible Taylor-Green vortex shows CPU and GPU dissipation histories agreeing to machine precision on LUMI and with published reference data. A single-node run on the ARCHER2 GPU platform provides an independent portability check. End-to-end performance is evaluated on fully populated nodes using Taylor-Green vortex, LM1021 sonic-boom, and NASA high-lift CRM benchmarks. Relative to the previous production implementation, the refactored CPU path is 1.27-1.67 times faster, while GPU offload delivers same-node speed-ups of 2.71-4.05 and 84-101 percent strong-scaling efficiency. The results show that standards-based OpenMP provides portable, production-scale acceleration without sacrificing numerical fidelity, CPU performance, or software sustainability.