arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2412.02175math.OCcs.LGstat.ML

光滑非凸优化的改进复杂度:一种采用拟牛顿方法的双层在线学习方法

Improved Complexity for Smooth Nonconvex Optimization: A Two-Level Online Learning Approach with Quasi-Newton Methods

  • The University of Texas at Austin(德克萨斯大学奥斯汀分校)

机构由 AI 辅助整理,请以论文原文为准。

Ruichen Jiang, Aryan Mokhtari, Francisco Patitucci

更新

AI总结:

针对仅用梯度信息求光滑非凸函数 $ε$-FOSP 的问题,论文将其转化为双层在线学习并提出乐观拟牛顿算法,在 $d=O(ε^{-1/2})$ 时改进复杂度,并首次证明拟牛顿法有望优于梯度下降型方法。

AI中文摘要:

我们研究在仅能获取梯度信息的情况下,寻找光滑函数的一个 $ε$-一阶驻点(FOSP)的问题。在假设目标函数的梯度和 Hessian 均为 Lipschitz 连续的条件下,该任务已知最优的梯度查询复杂度为 ${O}(ε^{-7/4})$。在本工作中,我们提出一种梯度复杂度为 ${O}(d^{1/4}ε^{-13/8})$ 的方法,其中 $d$ 为问题维度;当 $d = {O}(ε^{-1/2})$ 时,这会带来改进的复杂度。为实现这一结果,我们设计了一种底层涉及求解两个在线学习问题的优化算法。具体而言,我们首先将为非凸问题寻找驻点的任务,重新表述为在一个在线凸优化问题中最小化遗憾,其中损失由目标函数的梯度决定。随后,我们引入一种新颖的乐观拟牛顿方法来求解该在线学习问题,其中 Hessian 近似更新本身被构造为矩阵空间中的一个在线学习问题。除了改进利用梯度预言机获得 $ε$-FOSP 的复杂度界之外,我们的结果还首次给出保证,表明拟牛顿方法在非凸情形下有可能优于梯度下降型方法。

英文摘要:

We study the problem of finding an $ε$-first-order stationary point (FOSP) of a smooth function, given access only to gradient information. The best-known gradient query complexity for this task, assuming both the gradient and Hessian of the objective function are Lipschitz continuous, is ${O}(ε^{-7/4})$. In this work, we propose a method with a gradient complexity of ${O}(d^{1/4}ε^{-13/8})$, where $d$ is the problem dimension, leading to an improved complexity when $d = {O}(ε^{-1/2})$. To achieve this result, we design an optimization algorithm that, underneath, involves solving two online learning problems. Specifically, we first reformulate the task of finding a stationary point for a nonconvex problem as minimizing the regret in an online convex optimization problem, where the loss is determined by the gradient of the objective function. Then, we introduce a novel optimistic quasi-Newton method to solve this online learning problem, with the Hessian approximation update itself framed as an online learning problem in the space of matrices. Beyond improving the complexity bound for achieving an $ε$-FOSP using a gradient oracle, our result provides the first guarantee suggesting that quasi-Newton methods can potentially outperform gradient descent-type methods in nonconvex settings.

补充信息

↑