AI 中文总结
该研究为费米子VMC引入框架,用连续归一化流优化反对称基波函数,介绍多种架构,通过离线采样等技术减少计算时间和内存,在多GPU上实现强缩放性,训练出低基态能量,改进费米子基态估计方法。
AI 中文摘要
我们引入了一个用于费米子变分蒙特卡罗(VMC)的框架,其中连续归一化流(CNF)优化固定的反对称基波函数。该流实现为置换等变神经常微分方程,是一种学习基未捕捉到的相关性的光滑、拓扑保持映射;等变性保持基的反对称性,原则上该流可改进任何能有效采样的反对称假设。我们用斯莱特和贾斯特罗 - 斯莱特基证明了这一点,不过更具表现力的选择也是可行的。通过将预缓存的基样本推过正向常微分方程可获得流的玻恩分布的精确样本,训练时无需马尔可夫链蒙特卡罗(MCMC)。基样本离线生成并在多个训练批次和运行中复用,将样本生成与参数优化解耦,并实现跨多个GPU的并行训练。我们引入了三种新颖的置换等变向量场架构:成对深度集(PDS)、费米网向量场(FVF)和成对深度集梯度(PDSG),每种架构在表现力和计算成本上有不同平衡。我们还引入了用于动能计算的增强动力学公式,将所需导数量作为常微分方程状态变量共同演化,消除通过常微分方程轨迹的微分,显著减少了挂钟时间和内存。在谐波捕获的无自旋电子系统上的训练运行表明基态能量低于耦合簇单双激发(CISD)参考值。缩放实验表明,使用NERSC的Perlmutter超级计算机的32个GPU节点,对于三维中多达$N = 48$个粒子的系统,从1个NVIDIA A100扩展到128个时具有近乎理想的强缩放性。
英文摘要
We introduce a framework for fermionic variational Monte Carlo (VMC) in which a continuous normalizing flow (CNF) refines a fixed antisymmetric base wavefunction. The flow is implemented as a permutation-equivariant neural ODE, a smooth, topology-preserving map that learns correlations not captured by the base; equivariance preserves the antisymmetry of the base, so the flow can in principle improve any antisymmetric ansatz that can be sampled efficiently. We demonstrate this using Slater and Jastrow-Slater bases, though more expressive choices are admissible. Exact samples from the flow's Born distribution are obtained by pushing pre-cached base samples through the forward ODE, requiring no Markov chain Monte Carlo (MCMC) at training time. The base samples are generated offline and reused across training batches and runs, decoupling sample generation from parameter optimization and enabling embarrassingly parallel training across multiple GPUs. We introduce three novel permutation-equivariant vector field architectures: Pairwise Deep Sets (PDS), FermiNet Vector Fields (FVF), and Pairwise Deep Sets Gradient (PDSG), each offering a different balance of expressivity and computational cost. We further introduce an augmented dynamics formulation for kinetic energy computation that co-evolves the required derivative quantities as ODE state variables, eliminating differentiation through the ODE trajectory and yielding significant reductions in wall-clock time and memory. Training runs on systems of harmonically trapped spinless electrons demonstrate ground-state energies below CISD reference values. Scaling experiments demonstrate near-ideal strong scaling from 1 to 128 NVIDIA A100s using 32 GPU nodes of NERSC's Perlmutter supercomputer for systems of up to $N = 48$ particles in three dimensions.
Comments21 pages, 9 figures