arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TERRA-NG v1.0:极端规模、GPU加速的地幔对流

TERRA-NG v1.0: Extreme-Scale, GPU-accelerated Mantle Convection

Fabian Böhm, Nils Kohl, Ponsuganth Ilangovan, Gabriel Robl, Fatemeh Rezaei, Marcus Mohr, Bernhard S. A. Schuberth, Harald Köstler, Hans-Peter Bunge, Ulrich Rüde

arXiv 2609.21633首次发表:更新:

发表机构

Friedrich–Alexander–Universität Erlangen–Nürnberg; LMU Munich; CERFACS(弗里德里希·亚历山大大学埃尔朗根-纽伦堡分校; 慕尼黑大学; 欧洲流体动力学、空气动力学和燃烧研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

TERRA-NG是可移植、GPU加速的无矩阵地幔对流代码,通过单一Kokkos C++实现支持多厂商GPU,利用球形楔形网格优化,经基准验证,可在多台超级计算机上实现极端规模的高分辨率模拟。

AI 中文摘要

我们介绍TERRA-NG,一个可移植、GPU加速、无矩阵的地幔对流代码。单一的Kokkos C++实现可在NVIDIA、AMD和Intel GPU超级计算机上大规模运行。TERRA-NG的设计刻意精简:基于径向挤压的球形楔形网格,针对球形壳几何进行定制,从而实现了领域特定的优化,如单积分点积分评估、径向坐标存储压缩和径向共享内存分块。相应的低阶$W_1$-iso-$W_2/W_1$楔形基Stokes-能量离散化已通过Zhong等人(2008)的球形壳对流基准套件验证。我们通过JUWELS Booster(NVIDIA A100)、MareNostrum 5(NVIDIA H100)、LUMI-G(AMD MI250X)、Hunter(AMD MI300A APU)和SuperMUC-NG Phase 2(Intel PVC)超级计算机上的强扩展和弱扩展展示了TERRA-NG。在约11公里和约5.6公里径向间距(约28亿和约220亿自由度)下的耦合地幔对流模拟可在所有考虑系统的标准节点分区上常规运行。全球约每网格点1公里的地幔对流(约1.4万亿自由度)在极端规模分配下可行,而一次亚公里级英雄运行(约0.7公里网格间距,扩展至LUMI-G约11,000个GPU,约11万亿自由度)展示了该代码在未来更大机器上的潜力。

英文摘要

We present TERRA-NG, a portable, GPU-accelerated, matrix-free mantle-convection code. A single Kokkos C++ implementation runs at scale on NVIDIA, AMD, and Intel GPU supercomputers. TERRA-NG has a deliberately narrow design: built on a radially extruded mesh of spherical wedges, tailored to the spherical shell geometry, which enables domain-specific optimizations like single quadrature-point integral-evaluations, radial coordinate storage compression and radial shared-memory tiling. The corresponding low-order $W_1$-iso-$W_2/W_1$ wedge-based Stokes--energy discretisation is verified against the Zhong et al.(2008) spherical-shell convection benchmark suite. We showcase TERRA-NG through strong- and weak-scaling on the JUWELS Booster (NVIDIA A100), MareNostrum 5 (NVIDIA H100), LUMI-G (AMD MI250X), Hunter (AMD MI300A APU), and SuperMUC-NG Phase 2 (Intel PVC) supercomputers. Coupled mantle convection simulations at $\sim\!11$ km and $\sim\!5.6$ km radial spacing ($\sim 2.8$ B and $\sim 22$ B DoFs) can be run routinely on standard node partitions of all considered systems. Global $\sim\!1$ km-per-gridpoint mantle convection ($\sim 1.4$ T DoFs) is feasible on an extreme-scale allocation, and a sub-km hero-run at $\sim\!0.7$ km grid spacing scaling up to $\sim 11,000$ GPUs of LUMI-G ($\sim 11$ T DoFs) shows the potential of the code on future, larger machines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑