机器学习中自适应梯度优化器的控制理论框架
A Control Theoretic Framework for Adaptive Gradient Optimizers in Machine Learning
- Division of Data and Decision Sciences(数据与决策科学部)
- Tata Consultancy Services Research(塔塔咨询服务公司研究院)
- University of Maryland(马里兰大学)
- Department of Mechanical Engineering(机械工程系)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文基于状态空间建模与控制理论传递函数提出自适应梯度优化通用框架及 Adam 新变体 AdamSSM,证明其收敛性,并在图像分类和语言建模任务上验证其泛化与收敛优势。
AI中文摘要:
自适应梯度方法已在深度神经网络优化中得到广泛应用,近期例子包括 AdaGrad 和 Adam。尽管 Adam 通常收敛更快,但与经典随机梯度方法相比,Adam 的泛化能力较差,因此人们提出了 Adam 的变体,例如 AdaBelief 算法,以增强其泛化能力。本文构建了一个用于求解非凸优化问题的自适应梯度方法通用框架。我们首先在状态空间框架中对自适应梯度方法进行建模,从而能够为 AdaGrad、Adam 和 AdaBelief 等自适应优化器给出更简洁的收敛性证明。随后,我们利用经典控制理论中的传递函数范式,提出了 Adam 的一个新变体,并将其命名为 AdamSSM。我们在从平方梯度到二阶矩估计的传递函数中加入了一个适当的零极点对。我们证明了所提出的 AdamSSM 算法的收敛性。在使用 CNN 架构进行图像分类以及使用 LSTM 架构进行语言建模的基准机器学习任务上的应用表明,与近期的自适应梯度方法相比,AdamSSM 算法改善了泛化精度差距,并具有更快的收敛速度。
英文摘要:
Adaptive gradient methods have become popular in optimizing deep neural networks; recent examples include AdaGrad and Adam. Although Adam usually converges faster, variations of Adam, for instance, the AdaBelief algorithm, have been proposed to enhance Adam's poor generalization ability compared to the classical stochastic gradient method. This paper develops a generic framework for adaptive gradient methods that solve non-convex optimization problems. We first model the adaptive gradient methods in a state-space framework, which allows us to present simpler convergence proofs of adaptive optimizers such as AdaGrad, Adam, and AdaBelief. We then utilize the transfer function paradigm from classical control theory to propose a new variant of Adam, coined AdamSSM. We add an appropriate pole-zero pair in the transfer function from squared gradients to the second moment estimate. We prove the convergence of the proposed AdamSSM algorithm. Applications on benchmark machine learning tasks of image classification using CNN architectures and language modeling using LSTM architecture demonstrate that the AdamSSM algorithm improves the gap between generalization accuracy and faster convergence than the recent adaptive gradient methods.