AI 中文总结
该研究提出节点间断Galerkin方法,结合隐式扫描求解器与矩保持投影,在非结构网格上求解多维BGK方程,获高加速比并在超算上完成大自由度边值问题计算。
AI 中文摘要
我们提出了一种节点间断Galerkin方法,用于求解多维非结构网格上的Bhatnagar-Gross-Krook(BGK)动理学方程。该方法采用隐式、基于扫描的求解器,并结合麦克斯韦源的矩保持投影,以实现时间上的高阶精度,同时避免边界层及其他几何特征带来的限制性时间步长。我们通过将二维和三维非结构网格上的Sod激波问题与闭式解及高阶解对比,验证了该方法在连续介质极限下的正确性。对于三阶B稳定对角隐式Runge-Kutta方法,证明了其线性L2稳定性。该求解器采用基于空间域分解的混合并行方案,在CPU和GPU硬件上执行局部扫描。通过开发基于图的新GPU友好型扫描算法(使用Kokkos性能可移植性库实现),展示了平台可移植性;与64核AMD EPYC 9654 CPU相比,在NVIDIA H100 GPU上实现了超过20倍的加速。最后,我们在橡树岭领导计算设施的Frontier超级计算机上,针对一个具有2.77万亿相空间自由度的边值问题,展示了结果,该问题在1536个节点上使用6144个AMD MI250X GPU执行。
英文摘要
We present a nodal discontinuous Galerkin method for solving the Bhatnagar-Gross-Krook (BGK) kinetic equation on multi-dimensional, unstructured grids. The method uses implicit, sweep-based solvers and a moment-preserving projection of the Maxwellian source to enable high-order accuracy in time while avoiding restrictive time steps imposed by boundary layers and other geometry-induced features. We verify that the method is correct in the continuum limit by comparing to closed-form and high-order solutions of the Sod shock problem on 2 and 3D unstructured grids. Linear L2 stability is demonstrated for a B-stable diagonally implicit Runge-Kutta method of third order. The solver uses a hybrid parallel scheme based on spatial domain decomposition with local sweeps performed on CPU and GPU hardware. Platform-portability is demonstrated through the development of new GPU-friendly, graph-based sweep algorithms that are implemented using the Kokkos performance portability library and achieve greater than 20 times speedup on NVIDIA H100 GPUs compared to 64-core AMD EPYC 9654 CPUs. Finally, we show results on the Frontier supercomputer at the Oak Ridge Leadership Computing Facility for a boundary value problem with 2.77 trillion phase space degrees of freedom that executed on 1536 nodes utilizing 6144 AMD MI250X GPUs.
CommentsSubmitted to Journal of Computational Physics