发表机构
Institute of High Energy Physics, Chinese Academy of Sciences(中国科学院高能物理研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出基于 Kokkos 的性能可移植 Wilson 纯规范蒙特卡罗模拟实现,支持任意 SU(N) 和维度,多后端竞争原生性能,并在顶级超算上展示大规模加速。
AI 中文摘要
高性能计算系统的日益多样化使得针对特定架构分别实现格点规范场论算法变得维护成本高昂。我们提出 \texttt{kwqft},一个性能可移植的 Kokkos 实现,用于在任意时空维度下对 $\mathrm{SU}(N)$ 杨-米尔斯理论进行 Wilson 纯规范蒙特卡罗模拟。规范群阶数 $N$ 和维度 $D$ 是编译时参数。单一源代码面向 Serial、OpenMP、CUDA、HIP 和 SYCL 执行空间,MPI 边界交换与内部更新重叠。该实现重现了精确的二维 plaquette 值以及最高至 $\mathrm{SU}(17)$ 规范群的已发表三维和四维值。在 NVIDIA A100 上,Kokkos CUDA 后端与原生 CUDA 代码竞争力相当,SIMD 加速改善了 Armv9 处理器上的 OpenMP 路径,并在 LineShine 超级计算机(目前 TOP500 排名第一)上展示了 $\mathrm{SU}(4)$ 格点的大规模加速。
英文摘要
The increasing diversity of high performance computing systems makes separate, architecture specific implementations of lattice gauge theory algorithms costly to maintain. We present \texttt{kwqft}, a performance portable Kokkos implementation of Wilson pure gauge Monte Carlo simulation for $\mathrm{SU}(N)$ Yang-Mills theory in an arbitrary number of space-time dimensions. The gauge group order $N$ and the dimension $D$ are compile time parameters. A single source targets the Serial, OpenMP, CUDA, HIP, and SYCL execution spaces, with MPI halo exchange overlapped with interior updates. The implementation reproduces the exact two-dimensional plaquette and published three and four dimensional values for gauge groups up to $\mathrm{SU}(17)$. On an NVIDIA A100 the Kokkos CUDA backend is competitive with a native CUDA code, SIMD acceleration improves the OpenMP path on Armv9 processors, and a large scale speedup is demonstrated for an $\mathrm{SU}(4)$ lattice on the LineShine supercomputer, currently ranked first on the TOP500 list.
Comments14 pages, 5 tables, 4 figures