发表机构
KAIST; Natural Science Research Institute; Department of Mathematical Sciences(韩国科学技术院; 自然科学研究所; 数学科学系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究非凸优化中无调度方法,通过李雅普诺夫分析给出最坏情况收敛速率分析,还将无调度梯度下降建模为非自治动力系统,证明其在微小扰动下可避免严格鞍点,解释了该方法性能良好的原因。
AI 中文摘要
无调度方法因减轻学习率调度器设计和调整负担而备受关注,其性能有时优于有调度的优化器。尽管实证结果良好,但非凸优化中的收敛理论仍未充分探索。本文对标准形式的无调度梯度下降和无调度随机梯度下降进行最坏情况分析,基于李雅普诺夫分析表明它们达到一阶方法的最优最坏情况收敛速率,还证明了无调度梯度下降在微小一次性扰动下可避免严格鞍点,有助于更好理解其性能。
英文摘要
Schedule-Free methods have attracted growing interest for alleviating the burden of designing and tuning a learning rate scheduler, while matching and sometimes even outperforming optimizers with tuned schedulers. Despite their strong empirical results, their convergence theory in nonconvex optimization, where modern machine learning objectives typically arise, has remained largely unexplored. In this paper, we provide worst-case analyses of Schedule-Free gradient descent and Schedule-Free stochastic gradient descent, in their standard form and without auxiliary modifications or restrictive conditions, for smooth but possibly nonconvex objectives. Based on a Lyapunov analysis derived from the continuous-time limiting ordinary differential equation associated with these methods, we show that Schedule-Free gradient descent and Schedule-Free stochastic gradient descent achieve the optimal worst-case convergence rates attainable among first-order methods. We further formulate Schedule-Free gradient descent as a nonautonomous dynamical system and prove strict-saddle avoidance under an arbitrarily small one-time perturbation. These theoretical results provide a better understanding of the strong performance that Schedule-Free methods demonstrate.
Comments44+7 pages, 2 figures