发表机构
CISPA Helmholtz Center for Information Security; Technical University of Munich(CISPA赫尔姆霍茨信息安全中心; 慕尼黑工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究揭示大学习率下对角线性网络中漂移与增益的竞争,提出干预措施引导增益以恢复稀疏解,表明大学习率可控制隐式偏差。
AI 中文摘要
大学习率可以定性改变神经网络训练轨迹,通常将优化推向远离经典梯度流行为的区域。稳定性边缘(EoS)为理解此类学习率引发的动力学提供了有价值的视角。我们研究了对角线性网络中的相应动力学,发现两种不同的隐式偏差之间存在竞争,它们共同决定回归设置中恢复解的稀疏性。与捕捉梯度下降相对于梯度流累积的平均离散化误差的增益(Gain)互补,我们推导出一个密切相关但被忽视的量:漂移(Drift)。在大学习率下,它描述了不同离散化误差之间的不平衡,并代表优化轨迹中的系统性偏移。虽然增益在某些机制下单调增长,并可能偏向更密集、更平坦的插值器,但漂移的影响取决于其与潜在解的对齐,这可以抵消或增强增益的效果。因此,其行为驱动模型选择,尤其是在早期训练阶段。为了验证我们的理论见解,我们引入了一种干预措施,主动引导增益以恢复更尖锐、更稀疏的解。因此,我们的分析揭示了大学习率并不普遍阻碍稀疏解的恢复。相反,它们可以被利用来控制训练的隐式偏差。
英文摘要
Large learning rates can qualitatively change the trajectory of neural network training, often pushing optimization into regimes far from classical gradient-flow behavior. The Edge of Stability (EoS) offers a valuable lens on the dynamics such learning rates induce. We study corresponding dynamics in diagonal linear networks, where we uncover a competition between two distinct implicit biases that jointly determine the sparsity of the recovered solution in regression settings. Complementary to the Gain, which captures the average discretization error accumulated by Gradient Descent relative to Gradient Flow, we derive a closely associated but overlooked quantity: the Drift. Under large learning rates, it describes an imbalance between different discretization errors and represents a systematic shift in the optimization trajectory. While the Gain grows monotonically in certain regimes, and can bias towards denser, flatter interpolators, the impact of the Drift depends on its alignment with potential solutions, which can either counteract or reinforce the effect of the Gain. Consequently, its behavior drives model selection, particularly during early training epochs. To validate our theoretical insights, we introduce an intervention that actively steers the Gain to recover sharper, sparser solutions. Thus, our analysis reveals that large learning rates do not universally hinder the recovery of sparse solutions. On the contrary, they can be harnessed to control the implicit bias of training.