发表机构
University of Waterloo(滑铁卢大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文证明随机重配置中的对角位移充当统计滤波器,平衡表达性差距与噪声过拟合,并提出多位移SR(MS-SR)以降低验证风险和更新方差。
AI 中文摘要
随机重配置(SR)是神经量子态(NQS)的标准优化器,但现代NQS的参数数量往往远超蒙特卡洛样本数。我们表明,在此情形下,对角位移不仅仅是一个数值稳定器,它充当了有限样本泛化的统计滤波器。在固定波函数下,SR是从切向特征到中心化局部能量的岭回归。其残差是表达性差距,即虚时演化中位于当前切空间之外的部分。该差距在总体意义上与切空间正交,但有限批次使其表现为噪声,SR可能对此过拟合。因此,位移在收缩有用更新方向与拟合采样残差带来的方差之间取得平衡。在$4\ imes4$海森堡图上的精确诊断区分了过参数化的两种效应。更大的切空间在减少表达性差距时有帮助,但在过拟合固定差距时可能有害。在使用基础NQS训练的$100$位点横场伊辛族中,验证风险随位移呈U形变化而方差减小,与噪声岭模型一致。这一观点引出多位移SR(MS-SR),它在数据自适应位移处平均独立的岭求解,形成更丰富、方差更低的谱滤波器。检查点局部实验表明,与固定位移SR基线相比,MS-SR降低了验证风险和更新方差。我们进一步在配对在线训练延续中比较MS-SR和SR,并进行独立的端点能量评估和单独的更新成本基准测试。
英文摘要
Stochastic reconfiguration (SR) is the standard optimizer for neural quantum states (NQS), but modern NQS often have far more parameters than Monte Carlo samples. We show that in this regime the diagonal shift is more than a numerical stabilizer. It acts as a statistical filter for finite-sample generalization. At a fixed wave function, SR is ridge regression from tangent features to the centered local energy. Its residual is the expressivity gap, the part of imaginary-time evolution outside the current tangent space. This gap is orthogonal to the tangent space in population, but finite batches make it act as noise that SR can overfit. The shift therefore balances shrinkage of useful update directions against variance from fitting sampled residuals. Exact diagnostics on a $4\times4$ Heisenberg graph separate two effects of overparameterization. Larger tangent spaces help when they reduce the expressivity gap, but they can hurt when they overfit a fixed gap. In a $100$-site transverse-field Ising family trained with a foundation NQS, validation risk is U-shaped in the shift while variance decreases, matching the noisy-ridge model. This view leads to multi-shift SR (MS-SR), which averages independent ridge solves at data-adaptive shifts to form a richer, lower-variance spectral filter. Checkpoint-local experiments show that MS-SR lowers validation risk and update variance relative to the fixed-shift SR baseline. We further compare MS-SR and SR in paired online training continuations, with independent endpoint energy evaluations and a separate update-cost benchmark.
Comments25 pages, 7 figures