arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33193cs.LGquant-ph

表观压缩,真实稳定性:学习量子波函数的固有维度

Apparent Compression, Real Stability: The Intrinsic Dimension of Learning a Quantum Wavefunction

Lu Wei, Yufeng Wang, Chenfeng Cao, Haibin Ling

首次发表
浏览论文内容

中文总结 AI 辅助

本文测量变分蒙特卡洛中学习量子波函数的固有维度,发现小维度虽具误导性但带来真实稳定性,维度随相变上升且子空间训练避免发散。

中文摘要 AI 辅助

训练需要权重空间中的多少个方向?固有维度通过训练仍能达到目标精度的最小随机方向数来回答这个问题,而较小的数值已激发了诸如LoRA等参数高效方法。我们针对变分蒙特卡洛(VMC)方法进行测量,该方法训练一个神经网络来表示量子多体系统的基态。VMC是一个要求严苛的测试,因为网络自行生成训练样本且每个梯度都带有噪声,同时也是一个具有启发性的测试,因为精确答案已知且每次运行都可评分。我们仅训练一个小型潜向量,通过一个固定的随机映射将其转换为网络权重,且不改变标准的自然梯度优化器。我们发现,较小的维度可能具有误导性,而它带来的稳定性则是真实的。在一个具有硬符号模式的磁体上,无法表示符号的网络在28,642个方向中的8个方向上达到其最佳能量,但这仅仅是因为此类网络无法达到更低能量;一旦符号可学习,符号和幅度都不再廉价。维度在量子相变过程中上升,因此它以远低于拟合标度定律的计算量追踪态的难度,然而即使在基态近乎平凡的情况下,它也从未低于由随机子空间本身设定的下限。相比之下,在子空间中训练在我们的实验中从未发散,而相同设置下的全参数训练却发生了发散,一项使用匹配求解器的对照实验将这种差异归因于降维。

英文摘要

How many directions in weight space does training need? The intrinsic dimension answers this with the smallest number of random directions in which training still reaches a target accuracy, and small values have motivated parameter-efficient methods such as LoRA. We measure it for variational Monte Carlo (VMC), which trains a neural network to represent the ground state of a quantum many-body system. VMC is a demanding test, because the network generates its own training samples and every gradient is noisy, and a revealing one, because the exact answer is known and every run can be scored. We train only a small latent vector that a frozen random map turns into the network's weights, with no change to the standard natural-gradient optimizer. We find that a small dimension can be misleading, while the stability it brings is real. On a magnet with a hard sign pattern, a network that cannot represent signs reaches its best energy in 8 of 28,642 directions, but only because no such network can go lower; once signs are learnable, neither the signs nor the magnitudes are cheap. The dimension rises across a quantum phase transition, so it tracks how difficult a state is at far less compute than fitting a scaling law, yet it never falls below a floor set by the random subspace itself, even where the ground state is nearly trivial. Training in the subspace, in contrast, never diverged in our experiments, whereas full-parameter training with the same settings did, and a control with matched solvers attributes the difference to the reduced dimension.

补充信息

↑