arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越二次损失:Adam的稳定性相图

Beyond Quadratic Loss: The Stability Phase Diagram of Adam

Gaoxiang Tang, Huanran Chen, Ziming Liu

arXiv 2609.18314首次发表:更新:

发表机构

IIIS, Tsinghua University; College AI, Tsinghua University; Shanghai Qizhi Institute; MetaCircle(清华大学交叉信息研究院; 清华大学人工智能学院; 上海期智研究院; MetaCircle)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过映射Adam优化器的动量参数平面,发现损失尖峰由动量时间尺度失配和超二次损失几何共同决定,并建立了线性相边界理论。

AI 中文摘要

损失尖峰是神经网络训练中反复出现的不稳定性,可能由多种机制引起。对于Adam优化器而言,宏观损失尖峰与优化器动力学相关,但其两个动量时间尺度如何控制这些尖峰仍不清楚。我们通过映射$(\beta_1,\beta_2)$平面上的训练动力学来研究这种依赖性。在一系列模型-任务设置中,近似线性的边界$1-\beta_2=C(1-\beta_1)$将尖峰动力学与非尖峰动力学分开,而一维二次损失则产生近似三次方的斜率。一维超二次损失$L(x)\propto|x|^n$恢复了近线性缩放,并将边界系数与有效损失指数$n$联系起来。我们进一步表明,高置信度的交叉熵损失会形成一种核心-壁景观,包含一个狭窄的二次核心,随后是陡峭的壁,这在优化器更新的尺度上产生了有效的超二次行为。这些结果共同将Adam损失尖峰与动量时间尺度之间的不匹配以及超越Hessian的有限尺度超二次损失几何联系起来。

英文摘要

Loss spikes are recurrent instabilities in neural-network training and can arise from multiple mechanisms. For Adam in particular, macroscopic loss spikes have been linked to optimizer dynamics, yet how its two momentum timescales govern them remains unclear. We investigate this dependence by mapping training dynamics across the $(β_1,β_2)$ plane. Across a range of model--task settings, an approximately linear boundary, $1-β_2=C(1-β_1)$, separates spiky from non-spiky dynamics, whereas a one-dimensional quadratic loss produces approximately cubic slope. A one-dimensional superquadratic loss $L(x)\propto|x|^n$ recovers the near-linear scaling and links the boundary coefficient to the effective loss exponent $n$. We further show that confident cross-entropy losses develop a core--wall landscape comprising a narrow quadratic core followed by a steep wall, which produces effective superquadratic behavior at the scale of an optimizer update. Together, these results connect Adam loss spikes to both the mismatch between momentum timescales and finite-scale superquadratic loss geometry beyond the Hessian.

Comments21 pages, 9 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑