arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2511.06675math.OCcs.LG

Adam对称性定理:随机Adam优化器收敛性的刻画

Adam symmetry theorem: characterization of the convergence of the stochastic Adam optimizer

  • School of Data Science, The Chinese University of Hong Kong, Shenzhen (CUHK-Shenzhen)(数据科学学院,香港中文大学(深圳))
  • School of Data Science and School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen (CUHK-Shenzhen)(数据科学学院和人工智能学院,香港中文大学(深圳))

机构由 AI 辅助整理,请以论文原文为准。

Steffen Dereich, Thang Do, Arnulf Jentzen, Philippe von Wurstemberger

AI总结:

针对强凸随机优化问题,证明了Adam优化器的收敛速率,并通过Adam对称性定理揭示了其收敛到全局最优解当且仅当数据对称分布,否则不收敛。

AI中文摘要:

除了标准的随机梯度下降(SGD)方法外,Kingma & Ba(2014)提出的Adam优化器目前可能是人工智能(AI)系统中深度神经网络训练中最著名的优化方法。尽管Adam广受欢迎且成功,但即使在强凸随机优化问题(SOP)类中,为Adam提供严格的收敛性分析仍然是一个开放的研究问题。在本工作的主要结果之一中,我们针对强凸SOP类建立了Adam关于梯度步数(关于学习率大小的收敛速率为1/2)、小批量大小(关于小批量大小的收敛速率为1)以及Adam二阶矩参数大小(关于二阶矩参数到1的距离的收敛速率为1)的收敛速率。在本工作的另一个主要结果中,我们称之为Adam对称性定理,通过证明对于一类特殊的简单二次强凸SOP,当且仅当SOP中的随机变量(SOP中的数据)对称分布时,Adam随着梯度步数趋于无穷收敛到SOP的解(强凸目标函数的唯一极小值点),从而说明了所建立收敛速率的最优性。特别地,在SOP中的随机变量非对称分布的标准情况下,我们反驳了Adam随着步数趋于无穷收敛到SOP的极小值点。我们还通过若干数值模拟补充了收敛性分析和Adam对称性定理的结论,这些模拟表明了所建立收敛速率的精确性,并说明了Adam对称性定理所揭示现象的实际出现情况。

英文摘要:

Beside the standard stochastic gradient descent (SGD) method, the Adam optimizer due to Kingma & Ba (2014) is currently probably the best-known optimization method for the training of deep neural networks in artificial intelligence (AI) systems. Despite the popularity and the success of Adam it remains an \emph{open research problem} to provide a rigorous convergence analysis for Adam even for the class of strongly convex SOPs. In one of the main results of this work we establish convergence rates for Adam in terms of the number of gradient steps (convergence rate \nicefrac{1}{2} w.r.t. the size of the learning rate), the size of the mini-batches (convergence rate 1 w.r.t. the size of the mini-batches), and the size of the second moment parameter of Adam (convergence rate 1 w.r.t. the distance of the second moment parameter to 1) for the class of strongly convex SOPs. In a further main result of this work, which we refer to as \emph{Adam symmetry theorem}, we illustrate the optimality of the established convergence rates by proving for a special class of simple quadratic strongly convex SOPs that Adam converges as the number of gradient steps increases to infinity to the solution of the SOP (the unique minimizer of the strongly convex objective function) if and \emph{only} if the random variables in the SOP (the data in the SOP) are \emph{symmetrically distributed}. In particular, in the standard case where the random variables in the SOP are not symmetrically distributed we \emph{disprove} that Adam converges to the minimizer of the SOP as the number of Adam steps increases to infinity. We also complement the conclusions of our convergence analysis and the Adam symmetry theorem by several numerical simulations that indicate the sharpness of the established convergence rates and that illustrate the practical appearance of the phenomena revealed in the \emph{Adam symmetry theorem}.

补充信息

↑