发表机构
University of Münster; The Chinese University of Hong Kong, Shenzhen(明斯特大学; 香港中文大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对超参数完全可控的RMSprop优化器,解决了凸随机优化下误差常数一致有界的未决问题,给出了适用于所有梯度步的非渐近收敛速率估计,核心创新为RMSprop二阶矩过程的逆矩界。
AI 中文摘要
用于训练人工智能系统的流行自适应随机梯度下降(SGD)方法包括RMSprop、Adam和AdamW优化器,其中Adam和AdamW中的自适应部分基本与RMSprop一致。这类自适应方法涉及多个超参数,包括正则化参数ε(用于避免除以零,在PyTorch中通常被选为非常接近0的值,如默认值为10⁻⁸)和二阶矩衰减参数β(通常被选为非常接近1的值,如PyTorch中RMSprop的默认值为0.99,Adam和AdamW的默认值为0.999)。尽管这些方法相关性很高,但即使在凸随机优化问题的情况下,为这类方法提供误差估计且误差常数不会爆炸、而是相对于超参数一致有界,仍然是一个未解决的研究问题。本工作的核心贡献是针对RMSprop基本解决了该问题。具体而言,我们将RMSprop过程下目标函数停止评估的期望上界,由以下部分组成:随训练时间指数衰减的初始化项、阶为γₙ的随机近似余项,以及阶为(1-β)²的记忆误差,且误差常数在步长、二阶矩衰减参数β和正则化参数ε∈[0,1](也包含ε=0)的所有可允许选择上均被一致控制。我们的非渐近误差估计不仅对所有足够大的n成立,且对每一个梯度步n=1,2,3,…均成立,所有误差常数均被明确指定。我们分析证明中的关键创新新特征是RMSprop二阶矩过程的合适逆矩界。
英文摘要
Popular adaptive stochastic gradient descent (SGD) methods to train artificial intelligence (AI) systems include the RMSprop, the Adam, and the AdamW optimizers, where the adaptivity parts in Adam and AdamW basically just coincide with RMSprop. Such adaptive methods involve several hyperparameters including the regularization parameter $ε$ (which ensures that one does not divide by 0 and is often chosen to be very close to zero such as $10^{-8}$ in PyTorch by default) and the second moment decay parameter $β$ (which is often chosen to be very close to $1$ such as 0.99 (RMSprop) and 0.999 (Adam and AdamW) in PyTorch by default). Despite the high relevance of such methods, it remains an open research problem to provide error estimates for such methods with the error constants being not exploding but uniformly bounded with the respect to the hyperparameters, even in the situation of convex stochastic optimization problems. It is the key contribution of this work to essentially solve this problem for RMSprop. Specifically, we bound the expectation of the stopped evaluation of the objective function at the RMSprop process from above by the sum of an initialization term that decays exponentially in the training time, a stochastic approximation remainder of order $γ_n$, and a memory error of order $( 1 - β)^2$ with the error constants being uniformly controlled over all admissible choices of the step sizes, the second moment decay parameter $β$ and the regularization parameter $ε\in[0,1]$ (also covering $ε=0$). Our non-asymptotic error estimates hold not just for all sufficiently large n but hold for every gradient step $n=1,2,3,...$ with all error constants being explicitly specified. The key innovative new feature in the proof of our analysis are suitable inverse moment bounds for the second moment process in RMSprop.