arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

柱状等离子体的双分解多GPU粒子网格方法:批量傅里叶模式多重网格与平衡粒子板

Dual-decomposition multi-GPU particle-in-cell method for cylindrical plasmas: Batched Fourier-mode multigrid and balanced particle slabs

Yinjian Zhao, Xi Chen, Yingjie Chen

arXiv 2609.07172首次发表:更新:

发表机构

School of Energy Science and Engineering, Harbin Institute of Technology(哈尔滨工业大学能源学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出双分解多GPU粒子网格方法,通过傅里叶模式与粒子板分离及批量多重网格,实现高效求解与高扩展效率。

AI 中文摘要

三维静电粒子网格(PIC)模拟将不规则粒子操作与全局耦合的泊松求解相结合,而这两者的首选并行分解方式相互冲突。我们提出了一种双分解多GPU方法:在场求解过程中,完整的方位角傅里叶模式由单个GPU独占,而空间板则负责粒子。精确的全谱对角化和批量无矩阵几何多重网格将集合通信排除在V循环之外,之后每个GPU重建完整的场。对等迁移、单元重排序、warp聚合沉积以及容量受限的动态切分恢复了粒子的局部性和平衡性。CPU/GPU泊松解的相对L2误差低于6.9e-16;在八块V100 GPU上,所有65个物理模式在9.350毫秒内求解完成,相对误差低于8.2e-15。对于包含每物种707,788,800个粒子的相同512*128*400问题,完整的PIC循环每步耗时0.149108秒,从五块GPU扩展到八块GPU时强扩展效率为91.81%。一项独立的跨分辨率生产研究表明,相对于单块RTX 5090计算,聚合粒子更新和泊松单元吞吐量大约提高了三倍;这衡量的是细化问题的能力,而非硬件加速或形式收敛阶。

英文摘要

Three-dimensional electrostatic particle-in-cell (PIC) simulations combine irregular particle operations with a globally coupled Poisson solve, whose preferred parallel decompositions conflict. We introduce a dual-decomposition multi-GPU method: complete azimuthal Fourier modes are owned during the field solve, whereas spatial slabs own particles. Exact full-spectrum diagonalization and batched matrix-free geometric multigrid keep collectives outside the V-cycle, after which every GPU reconstructs the full field. Peer migration, cell reordering, warp-aggregated deposition, and capacity-constrained dynamic cuts restore particle locality and balance. CPU/GPU Poisson solutions agree to relative L2 errors below 6.9e-16; all 65 physical modes on eight V100 GPUs are solved in 9.350 ms with relative errors below 8.2e-15. For an identical 512*128*400 problem containing 707,788,800 particles per species, the complete PIC loop reaches 0.149108 s per step and strong-scales from five to eight GPUs with 91.81% efficiency. A separate cross-resolution production study shows approximately threefold aggregate particle-update and Poisson-cell throughput relative to a single RTX 5090 calculation; this measures refined-problem capability rather than a hardware speedup or formal convergence order.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑