AI 中文总结
该研究针对现有TN-VMC算法无法充分利用GPU加速的问题,开发适配GPU的向量化对称张量网络实现,在二维费米-哈伯德模型中实现费米子TN变分波函数最高300倍的GPU加速。
AI 中文摘要
基于张量网络(TN)的变分蒙特卡洛(VMC)计算近期在强关联自旋和费米子系统的基态计算中达到了颇具竞争力的精度,但现有张量网络变分蒙特卡洛(TN-VMC)算法的形式无法充分利用GPU加速。我们解决了关键缺失部分,即张量网络振幅与张量网络操作的向量化评估,通过对结构相同的计算进行批处理来确保高GPU利用率。具体而言,我们针对具有实际意义的阿贝尔对称张量网络(包含费米子张量网络),通过开发用于块稀疏张量表示与收缩的“扁平”张量网络形式主义实现向量化。在此基础上,我们构建了带有批处理张量网络计算的GPU适配对称TN-VMC工作流。在二维费米-哈伯德模型中,针对费米子TN变分波函数,我们展示了其相较于单核CPU实现最高达300倍的GPU加速效果。
英文摘要
Variational Monte Carlo (VMC) calculations based on tensor networks (TN) have recently achieved competitive accuracy in ground-state calculations of strongly correlated spin and fermionic systems. However, existing tensor network VMC (TN-VMC) algorithms have not been formulated in a manner that can fully utilize GPU acceleration. We tackle the key missing ingredient, namely, the vectorized evaluation of tensor network amplitudes and tensor network operations. This ensures high GPU utilization by batching over computations with identical structure. In particular, we show how to achieve vectorization for the practically relevant case of abelian symmetric tensor networks (which includes fermionic tensor networks) by developing a ``flat'' tensor network formalism for block-sparse tensor representation and contraction. Using this, we construct a GPU-adapted symmetric TN-VMC workflow with batched tensor network computation. In the two-dimensional Fermi--Hubbard model, we demonstrate a GPU speedup of up to $300 \times$ over single core CPU implementations, for fermionic TN variational wavefunctions.
Comments28 pages, 9 figures