arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2411.07061cs.LGmath.OCstat.ML

在线到非凸转换的通用框架:Schedule-free SGD 对非凸优化同样有效

General framework for online-to-nonconvex conversion: Schedule-free SGD is also effective for nonconvex optimization

Kwangjun Ahn, Gagik Magakyan, Ashok Cutkosky

首次发表 更新
浏览论文内容

中文总结 AI 辅助

本文提出了在线到非凸转换的通用框架,证明了 Schedule-free SGD 在非光滑非凸优化中具有最优迭代复杂度,并为其参数选择提供了理论见解。

中文摘要 AI 辅助

本研究受 Schedule-free 方法(由 A. Defazio 等人于 NeurIPS 2024 提出)在训练神经网络中显著的经验成功启发,探讨了该方法在非凸优化环境中的有效性。具体而言,我们证明了 Schedule-free SGD 在非光滑、非凸优化问题上实现了最优的迭代复杂度。我们的证明始于开发一个在线到非凸转换的通用框架,该框架将给定的在线学习算法转换为针对非凸损失的优化算法。该通用框架不仅涵盖了现有的转换方法,还催生了两种新颖的转换方案。值得注意的是,其中一种新转换直接对应于 Schedule-free SGD,从而使我们能够确立其最优性。此外,我们的分析为 Schedule-free SGD 的参数选择提供了有价值的见解,填补了凸理论无法解释的理论空白。

英文摘要

This work investigates the effectiveness of schedule-free methods, developed by A. Defazio et al. (NeurIPS 2024), in nonconvex optimization settings, inspired by their remarkable empirical success in training neural networks. Specifically, we show that schedule-free SGD achieves optimal iteration complexity for nonsmooth, nonconvex optimization problems. Our proof begins with the development of a general framework for online-to-nonconvex conversion, which converts a given online learning algorithm into an optimization algorithm for nonconvex losses. Our general framework not only recovers existing conversions but also leads to two novel conversion schemes. Notably, one of these new conversions corresponds directly to schedule-free SGD, allowing us to establish its optimality. Additionally, our analysis provides valuable insights into the parameter choices for schedule-free SGD, addressing a theoretical gap that the convex theory cannot explain.

发表机构

  • Microsoft Research(微软研究院)
  • MIT(麻省理工学院)
  • Boston University(波士顿大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑