发表机构
Center for Computational Science and Engineering, MIT; Department of Civil and Environmental Engineering, MIT; Institute for Data, Systems, and Society, MIT(麻省理工学院计算科学与工程中心; 麻省理工学院土木与环境工程系; 麻省理工学院数据、系统与社会研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究随机非凸优化能否黑盒简化为在线凸优化的静态遗憾最小化,提出维护梯度跟踪器并由在线学习器选预处理器的方法,对光滑和非光滑非凸目标均适用,解决了相关开放问题,还为理解自适应优化方法提供新视角。
AI 中文摘要
我们研究随机非凸优化是否能以黑盒方式简化为在线凸优化中的普通静态遗憾最小化。对于光滑非凸目标,我们的简化方法维护一个可预测梯度跟踪器,由黑盒在线学习器选择预处理器来确定如何将跟踪器转换为更新方向。学习器接收线性凸损失,并在一个无折扣在线游戏中与单个固定比较器进行评估。对于范围由\(M\)界定的\(\beta\)-光滑目标和方差由\(\sigma^2\)界定的无偏随机梯度预言机,我们建立了一个不等式。因此,任何具有\(\mathscr R_T(\mathcal A,I_d)=O(\sqrt T)\)的黑盒在线凸优化算法都能恢复经典的\(O(\frac{1}{\sqrt{T}})\)收敛率。我们还表明相同框架可扩展到非光滑的Lipschitz非凸目标。当在线凸优化预言机允许平方根静态遗憾时,转换对于相应的Goldstein驻点实现最优的\(O(T^{-2/7})\)收敛率。这些结果解决了Chen和Hazan(2024)提出的开放问题。更广泛地说,我们的框架将优化器设计分为梯度预测和在线预处理器选择,为理解自适应优化方法如AdaGrad和Shampoo提供了原则性视角,并可应用于非凸优化。
英文摘要
Stochastic nonconvex optimization is central to training deep networks and LLMs in modern machine learning. We give a black-box reduction from stochastic nonconvex optimization to ordinary static regret minimization in online convex optimization (OCO), thereby resolving the open problem posed by Chen and Hazan (2024). Our reduction maintains a predictable gradient tracker, while a black-box online learner $\mathcal{A}$ selects a preconditioner that transforms this tracker into the update direction. Given a \(β\)-smooth function with a range bounded by $M$ and an unbiased gradient oracle with variance bounded by $σ^2$, we bound the expected average squared gradient norm by $O(σ\sqrt{Mβ/T}+\sqrt{Mβ}\mathrm{Reg}_T(\mathcal{A})/T+\frac{Mβ}{T})$, where $\mathrm{Reg}_T(\mathcal{A})$ is the static regret of $\mathcal{A}$. Thus, any OCO oracle with $O(\sqrt{T})$ regret recovers the classical $O(T^{-1/2})$ convergence rate. We further extend the framework to nonsmooth nonconvex objectives, still relying only on ordinary static regret, and attain the optimal convergence rate for Goldstein-type stationarity. Finally, we conduct numerical experiments on nonconvex objectives to illustrate how the reduction exploits online-selected preconditioners while using the same stochastic-oracle budget as stochastic gradient descent.