arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从笔记本电脑到超级计算机具备可移植性的分钟级高保真陀螺动力学模拟

Minute-Scale High-Fidelity Gyrokinetic Simulations with Portability from Laptop to Supercomputer

Jian Bao, Huasheng Xie, Ming Yang, Zhixin Lu, Haotian Chen, Zhihong Lin, Feng Zhang

arXiv 2609.03354首次发表:更新:

发表机构

Beijing VeloAlpha Technology Co., Ltd.; Future Energy Institute, Shanghai Jiao Tong University; Max Planck Institut für Plasmaphysik(北京维奥阿尔法科技有限公司; 上海交通大学未来能源学院; 马克斯·普朗克等离子体物理研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究在PIF框架下提出混合谱方法并应用于GTC静电模型,可将粒子模拟加速48倍以上,单n案例加速超百倍,适用于从笔记本电脑到超级计算机的平台。

AI 中文摘要

全局陀螺动力学粒子模拟仍存在计算成本高的问题,因为其需要充足的粒子统计量和三维场求解器。本研究在粒子-傅里叶(PIF)框架内提出一种混合谱方法,并将其应用于GTC的静电模型中。电荷散射与场收集在二维极向网格上的粒子与场之间进行,而对应的泊松求解器采用径向有限差分和极向m谐波离散化。采用截断谱变换来连接场的多种表示,避免在粒子循环内对每个单独的m谐波执行代价高昂的粒子-网格操作。与传统粒子-网格(PIC)模拟的基准测试成功再现了单n离子温度梯度(ITG)模式结构和色散关系,以及多n非线性ITG输运及其由带状流的调控。与传统PIC相比,所提方法将有效问题规模减小了48倍以上,单n案例的加速比超过两个数量级。在笔记本电脑GPU上,约200万个粒子的2000步单n模拟耗时78.2秒,而多n湍流模拟也在数分钟内完成。此外,环形粒子移位通信的消除使其在多个NVIDIA A100 GPU上展现出良好的初步可扩展性。该数值方案广泛适用于加速从笔记本电脑到超级计算机平台的粒子模拟。

英文摘要

Global gyrokinetic particle simulations remain computationally expensive, as they demand both adequate marker statistics and three-dimensional field solvers. In this work, we present a hybrid spectral method within the particle-in-Fourier (PIF) framework and implement it in the electrostatic model of GTC. Charge scatter and field gather are performed between particles and fields on a two-dimensional poloidal mesh, while the corresponding Poisson solver is discretized using radial finite differences and poloidal $m$-harmonics. Truncated spectral transforms are employed to connect multiple representations for fields, avoiding costly particle-grid operations for each individual $m$-harmonic within the particle loop. Benchmarks against conventional particle-in-cell (PIC) simulations successfully reproduce single-$n$ ion temperature gradient (ITG) mode structures and dispersion relations, as well as multi-$n$ nonlinear ITG transport and its regulation by zonal flows. Compared to conventional PIC, the proposed method reduces the effective problem size by more than a factor of 48 and achieves a speedup of over two orders of magnitude for single-$n$ cases. A 2000-step single-$n$ simulation with approximately 2 million markers completes in 78.2 seconds on a laptop GPU, while multi-$n$ turbulence simulation also completes within minutes. Furthermore, the elimination of toroidal particle-shift communication yields promising preliminary scaling performance on multiple NVIDIA A100 GPUs. The numerical scheme is broadly applicable for accelerating particle simulations on platforms ranging from laptops to supercomputers.

Comments19 pages, 14 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑