通过更新的在线学习理解Adam优化器:Adam是伪装的FTRL
Understanding Adam Optimizer via Online Learning of Updates: Adam is FTRL in Disguise
- MIT(麻省理工学院)
- Microsoft Research(微软研究院)
- Georgia Tech(佐治亚理工学院)
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究基于在线学习视角揭示了Adam优化器等价于Follow-the-Regularized-Leader(FTRL)框架,从而阐明了其算法组件的重要性并提供了新的理论理解。
AI中文摘要:
尽管Adam优化器在实践中取得了成功,但对其算法组件的理论理解仍然有限。特别是,现有的大多数Adam分析所展示的收敛速率,SGD等非自适应算法也能轻易达到。在本研究中,我们基于在线学习提供了一个不同的视角,强调了Adam算法组件的重要性。受Cutkosky等人(2023)启发,我们考虑了一种称为更新/增量的在线学习框架,其中我们基于在线学习者来选择优化器的更新/增量。通过这一框架,设计一个好的优化器被简化为设计一个好的在线学习者。我们的主要观察是,Adam对应于一种称为Follow-the-Regularized-Leader(FTRL)的原则性在线学习框架。基于这一观察,我们从在线学习视角研究了其算法组件的优势。
英文摘要:
Despite the success of the Adam optimizer in practice, the theoretical understanding of its algorithmic components still remains limited. In particular, most existing analyses of Adam show the convergence rate that can be simply achieved by non-adative algorithms like SGD. In this work, we provide a different perspective based on online learning that underscores the importance of Adam's algorithmic components. Inspired by Cutkosky et al. (2023), we consider the framework called online learning of updates/increments, where we choose the updates/increments of an optimizer based on an online learner. With this framework, the design of a good optimizer is reduced to the design of a good online learner. Our main observation is that Adam corresponds to a principled online learning framework called Follow-the-Regularized-Leader (FTRL). Building on this observation, we study the benefits of its algorithmic components from the online learning perspective.