AI 中文总结
该研究将开源高阶CFD求解器HORSES3D移植到多GPU系统,验证了其正确性,在H100 GPU集群上实现高效扩展,可支撑工业级十亿自由度高阶CFD模拟。
AI 中文摘要
本文介绍了HORSES3D的GPU加速与大规模性能评估,HORSES3D是一款面向计算流体动力学的开源高阶间断伽略金求解器。该求解器通过OpenACC指令移植到NVIDIA GPU架构,在保留原始Fortran代码结构的同时,实现了主要计算内核的GPU驻留执行。该实现利用间断伽略金谱元法的单元局部结构,将单元级循环映射到GPU工作组,将节点操作映射到向量级并行。该GPU版本通过 manufactured解方法进行验证,并在经典湍流基准测试中得到确认。其性能在搭载NVIDIA H100 GPU的MareNostrum 5加速分区上进行评估,Taylor-Green涡基准测试显示,求解器效率随多项式阶数提升,且当每个GPU的工作量超过约16000至20000个单元时,可获得接近理想的强缩放和弱缩放性能。该求解器还针对高升力通用研究模型的翼身构型进行评估,该构型涉及复杂几何、真实边界条件及最多2080万个六面体单元的非结构网格。多项式阶数最高为P=7的模拟达到约107亿个自由度,且可高效扩展至2048个GPU。结果表明,HORSES3D在工业相关构型中保留了其性能特性,能够利用现代基于GPU的超级计算机开展十亿自由度级的高阶CFD模拟。
英文摘要
We present the GPU acceleration and large-scale performance assessment of HORSES3D, an open-source high-order discontinuous Galerkin solver for computational fluid dynamics. The solver is ported to NVIDIA GPU architectures using OpenACC directives, preserving the original Fortran code structure while enabling GPU-resident execution of the main computational kernels. The implementation exploits the element-local structure of discontinuous Galerkin spectral element methods by mapping element-level loops to GPU gangs and nodal operations to vector-level parallelism. The GPU version is verified using the method of manufactured solutions and validated on canonical turbulent-flow benchmarks. Its performance is assessed on the MareNostrum 5 accelerated partition using NVIDIA H100 GPUs. Taylor-Green vortex benchmarks show that solver efficiency improves with polynomial order and that near-ideal strong and weak scaling is obtained when the workload exceeds approximately 16,000 to 20,000 elements per GPU. The solver is further evaluated on the High-Lift Common Research Model wing-body configuration, which involves a complex geometry, realistic boundary conditions, and unstructured meshes with up to 20.8 million hexahedral elements. Simulations with polynomial orders up to $P=7$ reach approximately $10.7 \times 10^9$ degrees of freedom and scale efficiently to 2,048 GPUs. The results demonstrate that HORSES3D preserves its performance characteristics for industrially relevant configurations and can exploit modern GPU-based supercomputers for billion-degree-of-freedom high-order CFD simulations.