发表机构
Center for Computational Sciences, University of Tsukuba; Computing Research Center, High Energy Accelerator Research Organization (KEK)(筑波大学计算科学中心; 高能加速器研究机构(KEK)计算研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对少体玻色子系统,利用AI编码代理开发并优化了GPU并行SVM代码,引入久期方程法并批处理小矩阵,在GH200和MI300A上分别实现16.2倍和14.7倍加速,使6体及以上系统计算成为可能。
AI 中文摘要
随机变分方法(SVM)是在核物理和原子物理等多个领域中精确求解量子少体系统最强大的方法之一。据我们所知,目前尚未有报道针对GPU优化的SVM代码。我们开发了一个用于$^4$He原子少体团簇的并行SVM代码,并针对两种GPU架构(NVIDIA GH200和AMD Instinct MI300A)进行了优化,整个代码由AI编码代理编写。在算法方面,我们引入了结合Gu--Eisenstat方法的久期方程法,用于在SVM框架内求解广义特征值问题。代码调优的主要部分是将许多小矩阵分批处理,以便高效利用GPU,同时采用针对合并寻址优化的内存阵列布局。对于具有4800个SVM基态的五体系统,调优后的SVM代码在NVIDIA GH200(AMD MI300A)上的运行速度比在Intel Xeon Max CPU系统上运行相同调优代码快16.2(14.7)倍,比首个可工作的CPU实现快168(153)倍。所达到的性能使6体及更大团簇系统的研究成为可能。
英文摘要
The stochastic variational method (SVM) is one of the most powerful methods to solve quantum few-body systems precisely in various fields, such as nuclear and atomic physics. To the best of our knowledge, no SVM code optimized for GPUs has been reported. We developed a parallel SVM code for few-body clusters of $^4$He atoms and optimized it for two GPU architectures, NVIDIA GH200 and AMD Instinct MI300A, with the entire code written by an AI coding agent. On the algorithmic side, we introduce the secular-equation method with the Gu--Eisenstat prescription for solving the generalized eigenvalue problem within the SVM framework. The main part of the code tuning is to batch the many small matrices so that the GPUs are used efficiently, together with an array layout in memory optimized for coalesced addressing. For the five-body system with 4800 SVM basis states, the tuned SVM code runs 16.2 (14.7) times faster on the NVIDIA GH200 (AMD MI300A) than the same tuned code on the Intel Xeon Max CPU system, and 168 (153) times faster than the first working CPU implementation. The achieved performance brings 6-body and larger cluster systems within reach.
Comments7 pages, 3 figures, CANDAR-26 GCA workshop