发表机构
San Francisco University High School(旧金山大学附属高中)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文通过几何度量分析Adam与自然梯度下降的偏差,发现其优化能力源于结构近似误差与动量平滑的平衡,而非追踪自然梯度路径。
AI 中文摘要
Adam 是深度学习中的标准优化器,但其与自然梯度下降(NGD)的几何关系仍存在未解问题。我们研究了 Adam 的完整更新规则(包括动量),将其视为经过对角截断、经验标签替换和时间滞后处理的对角经验 Fisher 近似。利用尺度不变的 γ(Δθ) 度量,我们在四种损失景观下测量了 Adam 相对于真实 NGD 的几何偏差:良态线性回归、病态线性回归、逻辑回归以及一个非凸的小型神经网络。Adam 的几何轨迹具有上下文依赖性。在良态设置中偏差保持较低,但在病态条件下显著上升,在神经网络中达到约 10³ 的错位。较高的几何漂移与较慢的初始优化相关,但不会降低最终目标最小化效果;Adam 始终能达到较低的损失。此外,改进的经验 Fisher(iEF)比标准经验 Fisher(EF)追踪更稳定的路径,后者经常振荡或发散。我们的结果表明,Adam 的实际优化能力可能源于结构近似误差与动量平滑之间的平衡,而非对自然梯度路径的紧密追踪。
英文摘要
Adam is the standard optimizer in deep learning, yet its geometric relationship to natural gradient descent (NGD) contains unresolved questions. We study Adam's full update rule, including momentum, as a diagonal empirical Fisher approximation subject to diagonal truncation, empirical label substitution, and temporal lag. Using the scale-invariant $γ(Δθ)$ metric, we measure Adam's geometric deviation from true NGD across four loss landscapes: well-conditioned linear regression, ill-conditioned linear regression, logistic regression, and a non-convex small neural network. Adam's geometric trajectory is context-dependent. Deviation remains low in well-conditioned settings but rises significantly under ill-conditioning, reaching misalignments of $\approx 10^3$ in the neural network. Higher geometric drift correlates with slower initial optimization but does not degrade final objective minimization; Adam consistently reaches low loss. Furthermore, the improved empirical Fisher (iEF) tracks more stable paths than the standard empirical Fisher (EF), which frequently oscillates or diverges. Our results suggest Adam's practical optimization power may stem from a balance of structural approximation errors and momentum smoothing rather than close tracking of the natural gradient path.
Comments9 pages, 4 figures, 2 tables