arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习作为一种几何相变:深度网络中的重整化群流与各向异性对称性破缺

Learning as a Geometric Phase Transition: Renormalization Group Flow and Anisotropic Symmetry Breaking in Deep Networks

G. Le Pera, C. Nordio

arXiv 2607.10770首次发表:更新:

AI 中文总结

该研究将特征学习视为提升张量积学习度量的几何临界现象,推导相关展开与流,从微观运动学找\(\beta\)函数源,揭示威尔逊深度重整化群流受缺陷控制,还联系了时间随机训练动力学等,展现学习中几何相变等特性。

AI 中文摘要

我们将特征学习公式化为提升的张量积学习度量的几何临界现象。核心对象不是标量重叠,而是 \(\mathcal N_{0,L}=\frac1N\sum_{r=1}^{L}\Sigma_{r\to L}\otimes T_{0\to r - 1}\) 的目标-活性几何,它将前向回拉生存与后向推前可视性纠缠在一起。中性相是目标各向同性的:在限制到端点目标活性状态并进行迹归一化后,提升的度量与单位矩阵成比例。学习对应于这个目标各向同性不动点的不稳定性以及无迹目标对齐本征张量的出现。我们推导了局部各向异性插入的离散戴森展开及其连续的卡兰-西曼齐克流。关键的是,在构建完整的时间平均场理论之前,我们直接从微观运动学中识别出 \(\beta\) 函数的局部空间源:异步梯度更新产生同步度量应变,其目标活性对称无迹分量充当类似曲率的缺陷。然后,威尔逊深度重整化群流由这些缺陷的传输、平衡和粗粒无关性控制。在无标度计数假设下,重尾谱作为目标活性提升几何的谱出现,在匹配的回拉-推前扇区中有指数相加。最后,我们将这个深度重整化群图景与时间随机训练动力学以及学习通道在经验权重格拉姆矩阵上的运动印记联系起来。

英文摘要

We formulate feature learning as a geometric critical phenomenon of the lifted tensor-product learning metric. The central object is not a scalar overlap, but the target-active geometry of \[ \mathcal N_{0,L}=\frac1N\sum_{r=1}^{L}Σ_{r\to L}\otimes T_{0\to r-1}, \] which entangles forward pullback survival with backward push-forward visibility. The neutral phase is target-isotropic: after restriction to endpoint target-active states and trace normalization, the lifted metric is proportional to the identity. Learning corresponds to an instability of this target-isotropic fixed point and to the emergence of traceless target-aligned eigentensors. We derive discrete Dyson expansions for local anisotropic insertions and their continuous Callan--Symanzik flow. Crucially, before constructing the full temporal mean-field theory, we identify the local spatial source of the $β$-functions directly from microscopic kinematics: asynchronous gradient updates generate synchronous metric strains, whose target-active symmetric traceless components act as curvature-like defects. The Wilsonian depth RG flow is then governed by the transport, balance, and coarse-grained irrelevance of these defects. Heavy-tailed spectra arise, under a scale-free counting hypothesis, as the spectrum of the target-active lifted geometry, with exponent addition in the matched pullback--push-forward sector. Finally, we relate this depth RG picture to temporal stochastic training dynamics and to the kinematic imprint of the learned channel on empirical weight Gram matrices.

Comments11 pages. Working paper under technical review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑