发表机构
University of California, Berkeley(加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文介绍Chana,一种GPU加速的块结构自适应网格广义相对论磁流体动力学代码,集成多群M1中微子输运和Z4c动力时空演化,通过隐式处理辐射-物质相互作用,在A100 GPU上达到约1.1×10^6单元更新/秒,弱扩展效率超85%至256GPU。
AI 中文摘要
本文介绍了Chana,一种GPU加速的、块结构自适应网格代码,用于广义相对论磁流体动力学、多群M1中微子输运以及使用Z4c公式的动力时空演化。每个物理模块独立选择空间离散化和边界宽度,以平衡精度与计算成本。刚性辐射-物质相互作用采用隐式处理,其雅可比矩阵通过链式法则解析计算。该代码使用C++20编写,并利用MPI和Kokkos 5实现分布式内存和性能可移植的并行性。代码通过一系列GRMHD、辐射输运和动力时空测试得到验证。对于具有三种中微子种类、每类12个能量组的代表性核心坍缩超新星配置,该代码在NVIDIA A100 GPU上使用四阶RK4积分实现了约$1.1 \times 10^6$单元更新/GPU秒。在Perlmutter上,弱扩展效率在256个GPU时仍保持在85%以上。
英文摘要
This paper presents Chana, a GPU-accelerated, block-structured adaptive-mesh code for general relativistic magnetohydrodynamics, multigroup M1 neutrino transport, and dynamical spacetime evolution using the Z4c formulation. Each physical module uses independently selected spatial discretizations and ghost-zone widths to balance accuracy and computational cost. Stiff radiation-matter interactions are treated implicitly, with Jacobians calculated analytically from the chain-rules. The code is written in C++20 and it uses MPI and Kokkos 5 for distributed-memory and performance-portable parallelism. The code is validated through a number of GRMHD, radiation transport, and dynamic spacetime tests. For a representative core-collapse supernova configuration with three neutrino species and 12 energy groups per species, the code achieves approximately $1.1 \times 10^6$ cell updates per GPU second on NVIDIA A100 GPUs using four-stage RK4 integration. Weak-scaling efficiency remains above 85\% up to 256 GPUs on Perlmutter.
Comments16 pages, 10 figures. Submitted to ApJS