arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.13335cs.LGcond-mat.dis-nncond-mat.stat-mech

神经二次型:用于突发学习与标度律的统一最小模型

Neural Quadratic Forms: A Unified Minimal Model for Sudden Learning and Scaling Laws

Liu Ziyin, Yizhou Xu, Tomaso Poggio, Isaac Chuang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出神经二次型模型,统一描述感知机等多种架构的突发学习与标度律,通过对称性推导其训练动力学为洛特卡-沃尔泰拉方程,经数值验证符合相关行为。

中文摘要 AI 辅助

在平滑代价函数上通过梯度下降训练的神经网络,其学习过程却呈阶段性:代价函数在长平台期保持稳定,随后突然下降;与此同时,训练损失遵循平滑幂律。这两种行为的变体出现在微观结构差异极大的架构中,是少数相关集体变量的特征。我们证明,对称性确定了这些变量的形式:网络层是可互换单元的和,因此对单元进行重新标记不会改变其结构;结合平滑性以及单元梯度在原点处消失的条件,对称性强制训练初始时近零权重展开的通用主导形式为二次型$\text{Tr}[WW^{\top}A(x)]$,其中所有架构细节都被限定在单个“结构矩阵”$A(x)$中,我们可为每种架构计算该矩阵。感知机、注意力层、混合专家(Mixtures of Experts)和卷积在不同$A$下统一为同一模型。其训练动力学在“序参量”$M=WW^{\top}$上闭合,且当数据矩阵共享本征基时,可简化为洛特卡-沃尔泰拉(Lotka–Volterra)方程,其模式逐一激活。初始权重越小,激活时间间隔越大,平台期表现为平滑流的奇异极限;当多个模式未被解析时,相同事件会合并为训练时间的幂律,其指数可由该理论预测。我们在多种训练方法和架构上通过数值验证了这两种行为。

英文摘要

Neural networks trained by gradient descent on a smooth cost function can nevertheless learn in steps: the cost holds on long plateaus and then drops abruptly. Meanwhile, training losses instead follow smooth power laws. Variants of both behaviors occur in architectures with very different microscopic structures, which is the signature of a few relevant collective variables. We show that a symmetry fixes what those variables are: a network layer is a sum over interchangeable units, so relabeling the units leaves it unchanged; given smoothness and the condition that a unit's gradient vanish at the origin, symmetry then enforces a universal leading form for the expansion about the near-zero weights present at the start of training, the quadratic $\Tr[WW^{\top}A(x)]$, in which every architectural detail is confined to a single ``structure matrix" $A(x)$ that we compute for each architecture. Perceptrons, attention layers, mixtures of experts, and convolutions become one model at different $A$. Its training dynamics then close on the ``order parameter" $M=WW^{\top}$ and, whenever the data matrices share an eigenbasis, reduce to a Lotka--Volterra equation whose modes switch on one after another. The smaller the initial weights, the further apart the switch-on times, and the plateaus appear as a singular limit of a smooth flow; when many modes are unresolved the same events merge into a power law in training time whose exponent the theory predicts. We confirm both numerically across training methods and architectures.

发表机构

  • Massachusetts Institute of Technology(麻省理工学院)
  • École Polytechnique Fédérale de Lausanne(洛桑联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

↑