arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21366cs.DC

TokaGLINT:面向全三维托卡马克电磁模拟的可扩展GPU定制隐式求解器

TokaGLINT: A Scalable GPU-Tailored Implicit Solver for Full 3D Tokamak Electromagnetic Simulations

  • Computer Network Information Center, Chinese Academy of Sciences(中国科学院计算机网络信息中心)
  • University of Chinese Academy of Sciences(中国科学院大学)
  • Tsinghua University(清华大学)
  • University of Science and Technology of China(中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

Zifan Yang, Haoyuan Zhang, Jialin Li, Wu Yuan, Xiazhen Liu, Jian Zhang, Jianyuan Xiao, Shan Liang

AI总结:

提出面向全三维托卡马克电磁模拟的GPU加速隐式求解器TokaGLINT,通过层次化区域分解与张量求解器协同设计,在超1万GPU上实现90.1%弱扩展和53.9%强扩展效率,较基线加速2.67倍,并在EAST模拟中验证。

AI中文摘要:

我们介绍了TokaGLINT,一种用于全三维托卡马克模拟中电磁场计算的GPU加速隐式求解器,旨在实现高效的大规模并行GPU计算。其核心创新在于层次化区域分解与快速精确局部求解器的协同设计,其中层次化划分针对卡内细粒度子域进行定制,并利用基于张量的求解器来处理曲线坐标系辛CN-FDTD离散化的三维麦克斯韦方程组。在针对卡内多子域结构定制的自动算子融合与批处理支持下,该求解器通过离散变换解耦未知量,并利用张量结构化计算实现高硬件利用率,同时保持辛离散格式的长时间稳定性特征。TokaGLINT将电磁场求解规模扩展到超过10,000个GPU,实现了90.1%的弱扩展效率和53.9%的强扩展效率,同时相较于未预处理的BiCGStab基线(支持HIP的HYPRE)实现了2.67倍的单节点加速。该求解器已在SymPIC等离子体模拟代码中的EAST托卡马克模拟中得到验证,实现了高保真度的长时间建模。

英文摘要:

We introduce TokaGLINT, a GPU-accelerated implicit solver for electromagnetic field computations in full 3D tokamak simulations, aimed at efficient large-scale parallel GPU computing. Its central innovation lies in the co-design of hierarchical domain decomposition and a fast exact local solver, where hierarchical partitioning is tailored to match fine-grained intra-card subdomains and exploit the tensor-based solver dedicated to curvilinear-coordinate symplectic CN-FDTD-discretized 3D Maxwell equations. Backed by automated operator fusion and batching customized for the intra-card multi-subdomain structure, the solver decouples unknowns through discrete transformations and leverages tensor-structured computations to achieve high hardware utilization, while preserving the long-time stability characteristic of symplectic discretizations. TokaGLINT scales the electromagnetic field solve beyond 10,000 GPUs, achieving 90.1% weak and 53.9% strong scaling efficiency, while delivering a 2.67X single-node speedup over an unpreconditioned BiCGStab baseline (HIP-enabled HYPRE). It is validated in EAST tokamak simulations within the SymPIC plasma simulation code, enabling high-fidelity long-duration modeling.

补充信息

↑