AI 中文总结
本文针对用于训练大规模AI系统的动量SGD优化器开展严格误差分析,基于学习率、小批量规模及单点凸性常数建立其收敛速率。
AI 中文摘要
随机梯度下降(SGD)优化方案是人工智能(AI)系统中深度神经网络(DNN)优化的首选方法。实际应用中,人们通常不使用标准SGD,而是采用Adam、AdamW和MUON等标准SGD的加速型、自适应型和/或归一化变体来训练大规模AI系统。这些流行优化器的加速(高阶收敛速度)均依赖于动量SGD优化器。本工作对动量SGD优化器进行了严格的误差分析,特别地,我们基于学习率(步长)的大小、小批量的大小以及单点凸性常数的大小,建立了该动量优化器的收敛速率。
英文摘要
Stochastic gradient descent (SGD) optimization schemes are the methods of choice for the optimization of deep neural networks (DNNs) in artificial intelligence (AI) systems. Often not the standard SGD method is used but instead suitable accelerated, adaptive, and/or normalized variants of standard SGD such as Adam, AdamW, and MUON are employed to train large scale AI systems in practically relevant settings. The acceleration (higher order convergence speed) in all these popular optimizers relies on the momentum SGD optimizer. In this work we provide a rigorous error analysis for the momentum SGD optimizer. In particular, we establish convergence rates for the momentum optimizer in terms of the size of the learning rate (step size), the size of the mini-batch, and the size of the one-point convexity constant.
Comments36 pages