一种面向CPU/DCU异构集群的Fourier-Bessel粒子模拟的HIP兼容加速器后端
A HIP-Compatible Accelerator Backend for Fourier-Bessel Particle-in-Cell Simulations on CPU/DCU Heterogeneous Clusters
另 2 家 · 查看机构详情
- National Supercomputing Center in Zhengzhou, Zhengzhou University(郑州超算中心,郑州大学)
- School of Computer and Artificial Intelligence, Zhengzhou University(郑州大学计算机与人工智能学院)
- School of Communication and Artificial Intelligence, School of Integrated Circuits, Nanjing Institute of Technology(南京工程学院通信与人工智能学院、集成电路学院)
- School of Physics and Laboratory of Zhongyuan Light, Zhengzhou University(郑州大学物理学院及中原之光实验室)
- Laboratory for Advanced Computing and Intelligence Engineering, Wuxi(无锡先进计算与智能工程实验室)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对FBPIC依赖Numba CUDA无法在HIP环境运行的问题,开发了HIP兼容后端,在V100上实现1.32-1.54倍加速,并在DCU平台高效运行,多DCU强扩展达1.88倍。
中文摘要 AI 辅助
FBPIC(Fourier-Bessel粒子模拟)是一种用于相对论等离子体和加速器物理的高性能模拟代码。其原有的加速器后端依赖于Numba CUDA,这限制了它在使用HIP(异构计算可移植接口)编程环境的加速器(如DCU(深度计算单元)加速器)上的直接部署。在本工作中,我们开发了一个与HIP兼容的加速器后端,使FBPIC能够在DCU平台上高效运行,同时保留其Python用户界面和高层次模拟工作流。对于所评估的LWFA(激光尾场加速)工作负载,所提出的后端在NVIDIA V100 GPU上相比原始FBPIC实现实现了1.32-1.54倍的加速,并在DCU平台上实现了高效执行。我们还总结了将FBPIC移植到DCU平台的关键经验教训。多DCU实验在四个加速器上实现了1.88倍的强扩展加速,在约68%的弱扩展效率下,聚合吞吐量增加了2.72倍,通信分析表明节点间通信和同步是主要的可扩展性限制。除了FBPIC之外,所提出的方法为在异构加速器平台上移植和优化其他使用Python开发的科学计算应用提供了实用参考。
英文摘要
FBPIC (Fourier-Bessel particle-in-cell) is a high-performance simulation code for relativistic plasma and accelerator physics. Its original accelerator backend relies on Numba CUDA, which limits its direct deployment on accelerators using the HIP (Heterogeneous-Compute Interface for Portability) programming environment, such as DCU (Deep Computing Unit) accelerators. In this work, we develop an accelerator backend compatible with HIP that enables FBPIC to run efficiently on DCU platforms while preserving its Python user interface and high level simulation workflow. For the evaluated LWFA (laser-wakefield acceleration) workloads, the proposed backend achieves 1.32-1.54x speedups over the original FBPIC implementation on an NVIDIA V100 GPU and enables efficient execution on the DCU platform. We also summarize the key lessons learned from porting FBPIC to the DCU platform. Multi-DCU experiments achieve a 1.88x strong-scaling speedup on four accelerators and a 2.72x increase in aggregate throughput at approximately 68\% weak-scaling efficiency, with communication analysis identifying inter-node communication and synchronization as the main scalability limitations. Beyond FBPIC, the proposed approach provides a practical reference for porting and optimizing other scientific computing applications developed with Python on heterogeneous accelerator platforms.