发表机构
Purdue University(普渡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究揭示了归一化网络中学习率调度与权重衰减通过参数范数相互作用的精确离散时间定律,该定律由单一标量控制,并预测了不稳定性边界,为尺度不变优化提供了统一解释和精确控制方法。
AI 中文摘要
归一化使得神经网络的大部分区域在效果上具有尺度不变性,从而引入了一个隐藏的反馈回路,其中学习率调度和权重衰减通过参数范数相互作用,以控制优化器采取的有效步长。我们证明这种相互作用由一个精确的离散时间定律支配:一个标量量捕获了所有的调度和衰减强迫,而范数增长则引起相反的几何自淬灭效应。这产生了一个清晰的边界,将有效学习率区域干净地划分为收缩主导和扩张主导两种机制。为了理解潜在机制,我们提供了一个完全求解的归一化回归模型的精确分析,其中动力学简化为二维,并表明平衡点本质上是不稳定的,这意味着带有权重衰减的常数学习率无法稳定地维持内部平衡,而是产生由离散时间雅可比结构驱动的循环行为。我们进一步通过统一的齐次优化器框架将这一视角扩展到各种优化器,揭示了自淬灭强度的结构性二分法,为自适应方法在归一化下表现出系统性较弱的稳定性提供了第一性原理解释。在动力系统和神经网络(MLP、CNN、GPT2 / MNIST、CIFAR、wikiText、OpenWebText)中,预测的定律以高精度成立,并能够通过所识别的标量直接控制训练,性能在预测的边界处急剧达到峰值。总之,这些结果隔离了一个控制尺度不变优化的单一量,为现代深度学习中的训练动力学、优化器行为和调度设计提供了精确且可操作的视角。代码可在以下网址获取:此 https URL。
英文摘要
Normalization renders large parts of neural networks effectively scale invariant, inducing a hidden feedback loop in which learning-rate schedules and weight decay interact through the parameter norm to control the effective step taken by the optimizer. We show that this interaction is governed by an exact discrete-time law: a single scalar quantity captures all schedule and decay forcing, while norm growth induces an opposing geometric self-quenching effect. This yields a sharp boundary that cleanly separates contraction- and expansion-dominated effective learning rate regimes. To understand the underlying mechanism, we provide exact analysis of a fully solved normalized regression model where the dynamics reduce to two dimensions and show that the balance point is intrinsically unstable, implying that constant learning rate with weight decay cannot stably maintain an interior equilibrium and instead produces recurrent behavior driven by discrete-time Jacobian structure. We further extend this perspective across optimizers through unified homogeneous-optimizer framework that reveals a structural dichotomy in self-quenching strength, providing a first-principles explanation for why adaptive methods exhibit systematically weaker stabilization under normalization. Across dynamical systems and neural networks (MLP, CNN, GPT2 / MNIST, CIFAR, wikiText, OpenWebText), the predicted law holds with high precision and enables direct control of training via the identified scalar, with performance peaking sharply at the predicted boundary. Together, these results isolate a single governing quantity for scale-invariant optimization, providing a precise and actionable lens on training dynamics, optimizer behavior, and schedule design in modern deep learning. Code is available in https://github.com/shasanamin/normalized-optimization-dynamics.