arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

jQMC:一个基于JAX的用于GPU加速计算的从头算量子蒙特卡罗程序包

jQMC: A JAX-based ab initio quantum Monte Carlo package designed for GPU-accelerated computing

Kousuke Nakano, Michele Casula

arXiv 2607.13781首次发表:更新:

AI 中文总结

该研究介绍了基于Python的jQMC程序包,用于从头算QMC模拟。它实现了VMC和LRDMC算法,采用特定波函数,利用JAX进行自动微分等,支持多GPU并行化。经测试,在GPU上运行时比TurboRVB更快,不同算法有不同程度性能提升。

AI 中文摘要

我们展示了jQMC,这是一个基于Python的用于从头算量子蒙特卡罗(QMC)模拟的计算程序包,专为现代GPU加速计算环境设计……为评估此实现的实际性能,我们在NVIDIA GPU(A100和H100)上对jQMC性能进行了基准测试并分析了CUDA内核。对于此处分析的测试用例,系统大小达160个电子,当jQMC在GPU上运行时,其当前版本比实现相同算法和波函数的Fortran90代码TurboRVB更快。在运行时间方面,VMC的提升可达一个数量级,而LRDMC的提升较为适中。

英文摘要

We present jQMC, a Python-based computational package for {\it ab initio} Quantum Monte Carlo (QMC) simulations, designed for modern GPU-accelerated computing environments. jQMC implements two well-established QMC algorithms: Variational Monte Carlo (VMC) and the lattice-regularized variant of Diffusion Monte Carlo (LRDMC). The employed wave function is a Jastrow factor combined with the antisymmetrized geminal power with spin-singlet and spin-triplet pairings, which contains the single Slater determinant as its special lowest-rank case. The wave function can be initialized from external Hartree-Fock/Density Functional Theory calculations through the TREX-IO library (a common wave-function format across electronic-structure packages) and optimized by stochastic reconfiguration and linear-method energy minimization. One of the prominent features of jQMC is its use of JAX, which enables automatic differentiation for wave function optimization and atomic force calculations, and allows the main QMC algorithms to be Just-In-Time (JIT) compiled and portable across CPU and GPU. jQMC is vectorized over walkers at the top level of the QMC algorithms, providing efficient intra-GPU~(CPU) vectorization. The multi-GPU~(CPU) parallelization is also supported through MPI and JAX sharding. To assess the practical performance of this implementation, we benchmarked jQMC performance on NVIDIA GPUs (A100 and H100) and analyzed CUDA kernels. For the test cases analyzed here, with system sizes up to 160 electrons, the current version of jQMC is faster than TurboRVB, a Fortran90 code implementing the same algorithms and wave functions, once jQMC is run on GPUs. In terms of wall-time, the gain can reach an order of magnitude for VMC, while it is more moderate for LRDMC.

Comments68 pages, 11 figures, 5 tables,

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑