Adam优化器的收敛速率
Convergence rates for the Adam optimizer
- Institute for Mathematical Stochastics, University of Münster(明斯特大学数学随机学研究所)
- Shenzhen Research Institute of Big Data(深圳大数据研究院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文针对Adam优化器在强凸二次随机优化等一大类问题上建立了最优收敛速率,提出Adam向量场刻画其宏观行为,并证明Adam通常收敛到该向量场的零点而非目标函数临界点。
AI中文摘要:
随机梯度下降(SGD)优化方法如今是人工智能系统中训练深度神经网络(DNN)的首选方法。在实际相关的训练问题中,通常采用的优化方案并不是普通的原始标准SGD方法,而是适当加速和自适应的SGD优化方法。迄今为止,这类加速和自适应SGD优化方法中最受欢迎的变体或许是由Kingma和Ba在2014年提出的著名Adam优化器。尽管Adam优化器在实现中广受欢迎,但即使在目标函数(人们打算最小化的函数)为强凸的简单二次随机优化问题情形下,为Adam优化器提供收敛性分析仍然是一个悬而未决的研究问题。在本工作中,我们通过为一大类随机优化问题建立Adam优化器的最优收敛速率来解决这一问题,特别是覆盖了简单的二次随机优化问题。我们收敛性分析的关键要素是一个新的向量场函数,我们建议将其称为Adam向量场。这个Adam向量场准确描述了Adam优化过程的宏观行为,但不同于所考虑的随机优化问题的目标函数(我们打算最小化的函数)的负梯度。特别地,我们的收敛性分析揭示出,Adam优化器通常不会收敛到所考虑的优化问题的目标函数的临界点(目标函数梯度的零点),而是以一定速率收敛到这个Adam向量场的零点。
英文摘要:
Stochastic gradient descent (SGD) optimization methods are nowadays the method of choice for the training of deep neural networks (DNNs) in artificial intelligence systems. In practically relevant training problems, usually not the plain vanilla standard SGD method is the employed optimization scheme but instead suitably accelerated and adaptive SGD optimization methods are applied. As of today, maybe the most popular variant of such accelerated and adaptive SGD optimization methods is the famous Adam optimizer proposed by Kingma & Ba in 2014. Despite the popularity of the Adam optimizer in implementations, it remained an open problem of research to provide a convergence analysis for the Adam optimizer even in the situation of simple quadratic stochastic optimization problems where the objective function (the function one intends to minimize) is strongly convex. In this work we solve this problem by establishing optimal convergence rates for the Adam optimizer for a large class of stochastic optimization problems, in particular, covering simple quadratic stochastic optimization problems. The key ingredient of our convergence analysis is a new vector field function which we propose to refer to as the Adam vector field. This Adam vector field accurately describes the macroscopic behaviour of the Adam optimization process but differs from the negative gradient of the objective function (the function we intend to minimize) of the considered stochastic optimization problem. In particular, our convergence analysis reveals that the Adam optimizer does typically not converge to critical points of the objective function (zeros of the gradient of the objective function) of the considered optimization problem but converges with rates to zeros of this Adam vector field.