arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ANRe-M1:一种具有多能量M1中微子输运的GPU加速数值相对论代码

ANRe-M1: a GPU-accelerated numerical relativity code with multi-energy M1 neutrino transport

Takami Kuroda, Masaru Shibata

arXiv 2608.26681首次发表:更新:

AI 中文总结

本文提出ANRe-M1代码,其基于Kokkos实现,采用BSSN-Z4c形式化方法,经测试可高效运行于AMD MI300A集群,比原CPU版快约一个数量级,适用于全广义相对论下的多维超新星模拟。

AI 中文摘要

我们提出了ANRe-M1(Accelerated Numerical Relativity code with multi-energy M1 neutrino transport,即带多能量M1中微子输运的加速数值相对论代码),这是一种用于全广义相对论下多维核心坍缩超新星模拟的、性能可移植的GPU加速代码。ANRe-M1基于我们原有的基于CPU的Fortran框架构建,采用C++结合Kokkos实现。它保留了底层的广义相对论中微子辐射流体动力学算法,同时为适配加速器端执行重新设计了数据布局、并行分解和内存访问模式。时空部分采用BSSN-Z4c形式化方法,并与相对论流体动力学及依赖能量的M1中微子输运耦合。我们通过一系列流体动力学、辐射输运和动力学时空测试,以及与早期代码对比研究的结果进行比较,验证了该实现的正确性。此外,我们通过一项包含多能量中微子输运的代表性三维恒星坍缩计算,展示了其适用性。对于此处考虑的基准测试和计算系统,ANRe-M1在AMD MI300A加速处理单元集群上的运行速度,比CPU集群上的Fortran版本快约一个数量级。当分配的MI300A APU数量从2增加到128时,总吞吐量提升了39.7倍,相对于2个APU的情况,弱缩放效率为62%。这些结果确立了ANRe-M1作为一种高效且可移植的框架,可用于当前及未来基于加速器的超级计算机上的长期多维超新星模拟。

英文摘要

We present ANRe-M1 (Accelerated Numerical Relativity code with multi-energy M1 neutrino transport), a performance-portable, GPU-accelerated code for multidimensional core-collapse supernova simulations in full general relativity. ANRe-M1 builds on our original CPU-based Fortran framework and is implemented in C++ using Kokkos. It retains the underlying general-relativistic neutrino-radiation-hydrodynamics algorithms while redesigning the data layout, parallel decomposition, and memory-access patterns for accelerator-resident execution. The spacetime sector employs the BSSN--Z4c formulation and is coupled to relativistic hydrodynamics and energy-dependent M1 neutrino transport. We validate the implementation through a suite of hydrodynamic, radiation-transport, and dynamical-spacetime tests, together with comparisons to results from an earlier code-comparison study. We further demonstrate its applicability with a representative three-dimensional stellar-collapse computation including multi-energy neutrino transport. For the benchmark and computing systems considered here, ANRe-M1 runs approximately an order of magnitude faster on a cluster of AMD MI300A accelerated processing units than the Fortran version on a CPU cluster. When the allocation is increased from 2 to 128 MI300A APUs, the aggregate throughput rises by a factor of 39.7, corresponding to a weak-scaling efficiency of 62 per cent relative to the two-APU case. These results establish ANRe-M1 as an efficient and portable framework for long-term, multidimensional supernova simulations on current and forthcoming accelerator-based supercomputers.

Comments14 pagfes, 10 figures, submitted to MNRAS

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑