发表机构
University of Michigan; Michigan Institute for Computational Discovery and Engineering(密歇根大学; 密歇根计算发现与工程研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出理论框架,分析基于Transformer的神经量子态在上下文学习中的泛化性能,建立MSE泛化误差界,证明误差随示例数和深度反比下降,所需深度随系统大小线性增长,并推广至完整量子态,数值模拟验证理论。
AI 中文摘要
基于现代深度学习架构的神经量子态已成为量子多体系统的强大表示。特别是,基于Transformer的神经量子态提供了能够捕获长程相关性的表达模型,其经验泛化性能最近已被展示。然而,对其泛化行为的理论理解仍然很大程度上未被探索。在本文中,我们开发了一个理论框架,用于分析基于Transformer的神经量子态在上下文学习下的泛化性质。我们建立了以均方误差(MSE)表示的严格的推理时泛化误差界,表明逐点预测误差随上下文示例数量和Transformer深度的增加而反比减少。我们进一步表明,实现此保证所需的Transformer深度仅随系统大小线性扩展——即连续系统中的粒子数或离散系统中的qudit数。基于此结果,我们将分析扩展到以秩一密度算子形式表示的完整量子态,并在物理约束下推导出连续和离散域上的基于MSE的泛化界。最后,数值模拟证实了我们的理论分析。
英文摘要
Neural quantum states based on modern deep learning architectures have emerged as powerful representations for quantum many-body systems. In particular, Transformer-based neural quantum states provide expressive models capable of capturing long-range correlations, and their empirical generalization performance has recently been demonstrated. However, a theoretical understanding of their generalization behavior remains largely unexplored. In this paper, we develop a theoretical framework to analyze the generalization properties of Transformer-based neural quantum states under in-context learning. We establish a rigorous inference-time generalization error bound in terms of mean squared error (MSE), showing that the pointwise prediction error decreases inversely with both the number of in-context examples and the depth of the Transformer. We further show that the Transformer depth required to achieve this guarantee scales only linearly with the system size--namely, the number of particles in continuous systems or the number of qudits in discrete systems. Building on this result, we extend our analysis to full quantum states formulated as rank-one density operators, and derive MSE-based generalization bounds over both continuous and discrete domains under physical constraints. Finally, numerical simulations corroborate our theoretical analysis.