发表机构
Zuse Institute Berlin; Technical University of Berlin; IMDEA Software Institute(柏林齐美研究所; 柏林工业大学; IMDEA软件研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对关于ℓₚ范数的Hölder光滑凸优化问题,开发两类算法填补了梯度范数最小化的复杂度差距,解决了p>2等未决问题,实现近最优梯度或acular复杂度。
AI 中文摘要
最小化凸函数的梯度是优化与学习任务中的重要问题,梯度是可直接计算的近似平稳性证明,其最小化通常比函数值最小化能带来更强的结果。本研究针对关于ℓₚ范数(p≥1)的(L,κ)-Hölder光滑凸函数,研究其梯度范数最小化问题,开发出能实现该问题近最优梯度或acular复杂度的算法。在光滑情形下,本研究的结果解决了此前未解决的p>2的情况;对于Hölder光滑目标函数,本研究填补了整个p范围内的复杂度差距,包括据本领域所知的欧几里得情形下的差距。本研究提供两类算法:第一类具有简单迭代形式,推广了镜像对偶现象,利用带误差和不精确计算的算法的对偶行为;第二类利用以不同近似解为中心的累积正则化项,通过依次最小化这些正则化项来获得近最优速率。
英文摘要
Minimizing gradients of a convex function is an important problem across optimization and learning tasks. The gradient provides a directly computable certificate of approximate stationarity, and its minimization usually implies stronger results than those for minimization of function values. In this work, we study gradient-norm minimization for convex functions that are $(L,κ)$-Hölder smooth with respect to the $\ell_p$-norms, $p \geq 1$. We develop algorithms that achieve near-optimal gradient-oracle complexity for this problem. In the smooth case, our results resolve the previously open setting $p>2$. For Hölder-smooth objectives, we close the complexity gap throughout the full $p$-range, including to the best of our knowledge, a gap in the Euclidean case. We provide two families of algorithms: the first one comes with a simple iteration and generalizes a phenomenon known as mirror duality, exploiting dual behaviours of algorithms with errors and inexact computations. The second makes use of accumulating regularizers centered at different approximate solutions, which we sequentially minimize in order to provide our near-optimal rates.
Comments21 pages, 1 figure, 1 table