arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

线性上下文学习中淬火临界性的朗道理论

Landau theory of quenched criticality in linear in-context learning

Daesik Kim, Sumin Choi, Hyojae Jeon, Jung Hoon Han

arXiv 2608.28059首次发表:更新:

AI 中文总结

本研究将线性上下文学习的双下降奇点视为淬火无序系统的临界现象,构建朗道理论揭示其微观起源,验证预测并为相关统计物理理解奠定基础。

AI 中文摘要

上下文学习(ICL)允许预训练模型从提示中提供的示例推断新任务,而无需更新其参数。在ICL的线性模型中,当预训练样本数量与可学习参数数量相当的时候,预测误差会出现双下降奇点。我们将这种插值奇点表述为淬火无序系统的临界现象。通过比较同一线性ICL模型的退火和淬火描述,我们确定学习参数的样本间关联波动是奇异误差的微观起源。通过对重归一化脊参数ξ的腔自洽方程进行积分,构建了朗道势。(磁化)序参量的角色由ξ扮演,而裸脊参数λ成为其共轭磁场。归一化样本复杂度τ充当温度,双下降奇点出现在临界温度τ_c=1处。朗道磁化率恰好是预测误差波动贡献中发散的量。序参量与无脊极限下经验松弛矩阵的零特征值比例密切相关,这些零特征值定义了学习动力学中的平坦方向。朗道理论通常是序参量的三次型,临界指数为(β_cr,δ_cr,γ_cr)=(1,2,1)。在大上下文 regime 中,会出现由序参量受抑制所表征的类赝能隙 regime。通过原始学习问题的数值解,独立验证了朗道理论的预测,具有良好的定量一致性。我们的结果为线性上下文学习中插值临界性的可靠统计物理理解铺平了道路。

英文摘要

In-context learning (ICL) allows a pretrained model to infer a new task from examples supplied in its prompt without updating its parameters. In linear models of ICL, the prediction error develops a double-descent singularity when the number of pretraining samples becomes comparable to the number of learnable parameters. We formulate this interpolation singularity as a critical phenomenon of a quenched disordered system. By comparing annealed and quenched descriptions of the same linear ICL model, we identify the connected sample-to-sample fluctuations of the learned parameters as the microscopic origin of the singular error. A Landau potential is constructed by integrating the cavity self-consistency equation for the renormalized ridge parameter $ξ$. The role of (magnetization) order parameter is played by $ξ$, while the bare ridge parameter $λ$ becomes its conjugate magnetic field. The normalized sample complexity $τ$ acts as a temperature and the double-descent singularity occurs at the critical temperature $τ_c =1$. The Landau susceptibility is precisely the quantity that diverges in the fluctuation contribution to the prediction error. The order parameter is closely related to the fraction of zero eigenvalues of the empirical relaxation matrix in the ridgeless limit, which define flat directions in the learning dynamics. The Landau theory is generically cubic in the order parameter with critical exponents $(β_{\rm cr},δ_{\rm cr},γ_{\rm cr})=(1,2,1)$. In the large-context regime, there appears a pseudogap-like regime characterized by suppressed order parameter. Predictions of the Landau theory are independently confirmed from numerical solutions of the original learning problem with good quantitative agreement. Our results pave the way for solid statistical-physics understanding of the interpolation criticality in linear in-context learning.

Comments17 pages, 7 figures (counting subfigures)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑