发表机构
Variational AI; The Voleon Group; Simon Fraser University(Variational AI; Voleon集团; 西蒙弗雷泽大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出函数空间中的统计力学框架,通过学习算子与态密度分析深度网络误差动力学,揭示低曲率方向偏好,为集体组织提供宏观视角。
AI 中文摘要
深度神经网络在巨大参数空间中尽管具有高度非线性动力学,却表现出规则的宏观行为。我们直接在函数空间中发展了一种学习的统计力学描述,将参数配置视为微观实现,将函数及其动力学算子视为宏观变量。对于均方损失,精确的误差动力学由学习算子 \\(M=JJ^\\) 控制。将条件随机动力学的动力学玻尔兹曼权重与参数空间态密度(其局部曲率定义了一个统计算子 \\(B\\))相结合,并对局部涨落积分,得到 $$ \Phi_{\mathrm{fluc}}(M;B)=\frac{\sigma_\xi^2}{2}\log\det(M^{-1}+B)+\mathrm{const}. $$ 在固定谱下,当 \\([M,B]=0\\) 时,该项在旋转下平稳,通过将 \\(M\\) 的大特征值与 \\(B\\) 的小特征值配对来最小化,并产生一个针对旋转失配的局部恢复贡献。对于在温和稳定统计条件下的 ReLU 型函数空间,\\(B=\sigma_\xi^2L^\mathcal K L\\),其中 \\(L\\) 度量粗粒化的二阶结构。因此,低 \\(B\\) 扇区对应于(在 \\(\mathcal K\\) 的有界各向异性范围内)低结构曲率,意味着沿平滑、数据自适应方向更快弛豫的偏好。这些结果将函数空间确定为研究学习中稳定集体组织的自然宏观层面。
英文摘要
In the kernel regime, neural-network learning inherits its preferences from a frozen spectrum. During feature learning, this spectrum evolves, yet networks retain systematic biases toward simple, smooth directions. We develop a function-space statistical framework explaining the origin of these preferences, treating functions and their learning operators as macroscopic variables, with parameterization entering through the multiplicity of parameter configurations realizing each function. For mean-squared loss, error relaxes exactly under the evolving learning operator $M=JJ^\ast$. Training stochasticity induces a Gaussian weight over function-space states, while parameter multiplicity contributes an entropic operator $B$, defined by the curvature of its log multiplicity. A local Laplace expansion yields the fluctuation free energy $Φ_{\mathrm{fluc}}(M;B)=\frac{σ_ξ^2}{2}\log\det(M^{-1}+B)+\mathrm{const}$, analogous to an Occam factor. Under mild statistical conditions, this free energy is rotationally stationary exactly when $[M,B]=0$, is minimized by pairing large eigenvalues of $M$ with small eigenvalues of $B$, and generates a local restoring force against mismatch. Learning is therefore biased toward faster relaxation along entropically cheaper directions. This preference strengthens with training noise and vanishes in the deterministic limit, beyond gradient-flow accounts of operator alignment. For ReLU networks, we relate entropic curvature to the minimal rearrangement of activation boundaries required for a functional change and bound this structural cost by directional smoothness. Consequently, smooth directions are preferentially learned faster, in a data-adaptive manner, even as the learning operator evolves.