arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RAMSES-GPU:图形处理器上具有磁流体动力学和自引力的逐单元自适应网格细化

RAMSES-GPU: Cell-by-Cell Adaptive Mesh Refinement with Magneto-Hydrodynamics and Self-Gravity on Graphics Processing Units

Romain Teyssier, Eric Moseley, Robert V. Caddy, Troels Haugbolle, James Sunseri, Jun-Young Lee, Jonathan Palafoutas, Yue Pan, William Groger, Tom Abel, Burlen Loring, Guillaume Barnier

arXiv 2607.20704首次发表:更新:

AI 中文总结

该研究在GPU上实现并优化宇宙学模拟代码RAMSES,介绍其管理AMR策略及跨硬件架构的可移植性,阐述流体动力学、粒子动力学和自引力处理方式,通过基准测试表明相比多核CPU有显著加速,还讨论了其他加速物理模块。

AI 中文摘要

我们展示了使用CUDA Fortran在图形处理器(GPU)上对宇宙学模拟代码RAMSES的实现和优化。这个加速版本将包括流体动力学、粒子动力学和自引力在内的主要计算例程移植到多GPU架构上。我们详细介绍了在GPU上管理逐单元自适应网格细化(AMR)的策略,利用桶排序和前缀和进行AMR级别排序,通过CUB库进行基数排序以实现希尔伯特键排序,以及使用线性探测的fnv64哈希表进行快速空间索引。通过调度器和C - Fortran包装器实现跨不同硬件架构的可移植性,直接调用从CUDA Fortran框架直接翻译的CUDA、HIP和Metal内核。流体动力学更新通过三层共享内存内核架构(名为石头、剪刀、布)管理的Godunov MUSCL - Hancock HLLC黎曼求解器执行。粒子质量沉积使用通过原子加法或前缀和优化的云内单元(CIC)插值,并结合踢 - 漂移 - 踢时间积分推进器。自引力通过在各个级别执行分层V循环的多重网格(MG)泊松求解器处理。在NVIDIA A100和H200 GPU上进行的性能基准测试表明,与多核CPU相比有显著加速,对于诸如Sedov爆轰波、分子核心坍缩和宇宙学模拟等标准测试问题,加速倍数达到10倍至100倍。最后,我们简要讨论了其他加速物理模块,包括平衡冷却、多方状态方程、理想和非理想磁流体动力学(MHD)以及恒星反馈。

英文摘要

We present the implementation and optimization of the cosmological simulation code RAMSES on Graphics Processing Units (GPUs) using CUDA Fortran. This accelerated version ports the main computational routines, including hydrodynamics, particle dynamics, and self-gravity, to multi-GPU architectures. We detail our strategy for managing cell-by-cell Adaptive Mesh Refinement (AMR) on the GPU, utilizing bucket sort with prefix sums for AMR level sorting, radix sort via the CUB library for Hilbert key ordering, and an fnv64 hash table with linear probing for fast spatial indexing. Portability across diverse hardware architectures is achieved via a dispatcher and C-Fortran wrappers, calling CUDA, HIP, and Metal kernels directly translated from the CUDA Fortran framework. Hydrodynamics updates are executed via a Godunov MUSCL-Hancock HLLC Riemann solver managed through a three-tier shared-memory kernel architecture (named rock, paper, and scissor). Particle mass deposition uses Cloud-in-Cell (CIC) interpolation optimized with atomic additions or prefix sums, combined with a kick-drift-kick time integration pusher. Self-gravity is handled via a Multigrid (MG) Poisson solver performing hierarchical V-cycles on individual levels. Performance benchmarks conducted on NVIDIA A100 and H200 GPUs demonstrate substantial accelerations compared to multi-core CPUs, yielding 10x up to a 100x speedup for standard test problems such as the Sedov blast wave, molecular core collapse, and cosmological simulations. Finally, we briefly discuss additional accelerated physics modules, including equilibrium cooling, polytropic equations of state, ideal and non-ideal magneto-hydrodynamics (MHD), and stellar feedback.

Comments9 pages, 4 figures, submitted to RASTI

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑