MADA:通过超梯度下降实现的元自适应优化器
MADA: Meta-Adaptive Optimizers through hyper-gradient Descent
- Amazon Web Services(亚马逊网络服务)
- Technion – Israel Institute of Technology(以色列理工学院)
- University of Minnesota(明尼苏达大学)
- Ecole Polytechnique Fédérale de Lausanne(洛桑联邦理工学院)
- University of California Los Angeles(加利福尼亚大学洛杉矶分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出统一元自适应优化器框架 MADA,通过超梯度下降在参数化优化器空间中动态选择,并引入 AVGrad;实验显示其在视觉、语言及 GPT-2 任务中稳定优于 Adam。
AI中文摘要:
自 Adam 推出以来,已有多种用于深度学习的新型自适应优化器被提出。这些优化器通常在某些任务上表现出色,但未必能在所有任务上一致地超越 Adam。在本工作中,我们提出了元自适应优化器(Meta-Adaptive Optimizers,MADA),这是一个统一的优化器框架,能够泛化若干已知优化器,并在训练过程中动态学习最合适的优化器。MADA 的核心思想是对优化器空间进行参数化,并在训练期间使用超梯度下降在该空间中动态搜索。我们在视觉和语言任务上将 MADA 与其他流行优化器进行了实证比较,发现 MADA 始终优于 Adam 及其他流行优化器,并且对次优调节的超参数具有鲁棒性。在 GPT-2 训练和微调期间,与其他流行优化器相比,MADA 相对于 Adam 取得了更大的验证性能提升。我们还提出了 AVGrad,这是对 AMSGrad 的一种修改,用平均操作替代最大值操作,更适合超梯度优化。最后,我们给出了收敛分析,表明优化器的参数化插值可以改善其误差界(至多相差常数),这暗示了元优化器的优势。
英文摘要:
Following the introduction of Adam, several novel adaptive optimizers for deep learning have been proposed. These optimizers typically excel in some tasks but may not outperform Adam uniformly across all tasks. In this work, we introduce Meta-Adaptive Optimizers (MADA), a unified optimizer framework that can generalize several known optimizers and dynamically learn the most suitable one during training. The key idea in MADA is to parameterize the space of optimizers and dynamically search through it using hyper-gradient descent during training. We empirically compare MADA to other popular optimizers on vision and language tasks, and find that MADA consistently outperforms Adam and other popular optimizers, and is robust against sub-optimally tuned hyper-parameters. MADA achieves a greater validation performance improvement over Adam compared to other popular optimizers during GPT-2 training and fine-tuning. We also propose AVGrad, a modification of AMSGrad that replaces the maximum operator with averaging, which is more suitable for hyper-gradient optimization. Finally, we provide a convergence analysis to show that parameterized interpolations of optimizers can improve their error bounds (up to constants), hinting at an advantage for meta-optimizers.