发表机构
The University of Tokyo; RIKEN(东京大学; 理化学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出自适应双预处理梯度下降(adaptive DPGD),结合曲率预处理与无线性搜索步长规则,在凸和弱凸目标上实现线性收敛或渐近平稳,并通过半自动几何设计提升性能。
AI 中文摘要
非线性预处理使得梯度更新能够适应目标函数曲率的增长,但将其设计与合适的步长选择相结合具有挑战性。我们提出了自适应双预处理梯度下降(adaptive DPGD),该方法将基于曲率的预处理器设计与基于局部信息的无线性搜索步长规则相结合。对于凸目标函数,我们在目标函数局部强凸性下建立了线性收敛性。通过调整步长规则,我们将该方法扩展到弱凸目标函数,并在没有全局Lipschitz光滑性的情况下建立了渐近平稳性。此外,我们的局部假设为预处理器设计提供了更大的自由度。我们利用这一自由度开发了一种由目标函数曲率引导的半自动配方,产生了一族预处理器,其中包括归一化梯度下降和双曲梯度下降所基于的预处理器。在真实数据上的凸和弱凸问题的实验证明了预处理器设计和自适应步长选择的有效性。
英文摘要
Nonlinear preconditioning makes it possible to adapt gradient updates to the growth of the objective's curvature, but combining its design with suitable stepsize selection is challenging. We propose adaptive dual preconditioned gradient descent (adaptive DPGD), which combines curvature-based preconditioner design with a linesearch-free stepsize rule based on local information. For convex objectives, we establish linear convergence under local strong convexity of the objective. By adapting the stepsize rule, we extend the method to weakly convex objectives and establish asymptotic stationarity without global Lipschitz smoothness. Moreover, our local assumptions give greater freedom in preconditioner design. We exploit this freedom to develop a semiautomatic recipe guided by the objective's curvature, yielding a family of preconditioners that includes those underlying normalized and hyperbolic gradient descent. Experiments on convex and weakly convex problems with real data demonstrate the effectiveness of the preconditioner design and adaptive stepsize selection.
Comments42 pages, 16 figures