AI 中文总结
针对需不精确梯度的凸(L0, L1)-光滑优化问题,提出两种基于比较预言机的归一化一阶方法,推导收敛速率并通过实验验证理论结果。
AI 中文摘要
广义光滑性(如(L0, L1)-光滑性)近来备受关注,因其可建模现代机器学习与深度学习中出现的优化问题,而这类问题常违反梯度的经典Lipschitz假设。同时,在诸多应用中计算精确梯度可能不现实或计算成本过高。本研究仅基于近期提出的Comparison Oracle(比较预言机)开展凸(L0, L1)-光滑优化研究,该预言机可在线性时间内返回带界绝对误差的不精确归一化梯度。在此框架下,我们开发了归一化梯度下降(Normalized Gradient Descent)和带Polyak步长的梯度下降(Gradient Descent with Polyak stepsizes)的比较预言机变体,建立了保证收敛的近似误差显式上界,并推导了所有提出方法的收敛速率。与现有分析不同,我们的结果既不要求经典光滑性假设,也无需访问精确梯度或其精确归一化对应项。最后,数值实验证实了理论发现。
英文摘要
Generalized smoothness, such as (L0, L1)-smoothness, have recently attracted considerable attention due to their ability to model optimization problems arising in modern machine and deep learning, where the classical Lipschitz assumptions of the gradient is often violated. At the same time, computing exact gradients may be impractical or computationally expensive in many applications. In this work, we study convex (L0, L1)-smooth optimization (for normalized gradient method we consider quasi-convex problems too) under access only to a normalized approximation recently proposed Comparison Oracle, which returns an inexact normalized gradient in linear time with a bounded absolute error. Within this framework, we develop comparison-oracle variants of Normalized Gradient Descent and Gradient Descent with Polyak stepsizes. We establish explicit upper bounds on the approximation error that guarantee convergence and derive convergence rates for all proposed methods. Unlike existing analyses, our results require neither classical smoothness assumptions nor access to exact gradients or their exact normalized counterparts. Finally, numerical experiments corroborate the theoretical findings.
CommentsGeneralized Smoothness, Comparison Oracle, Inexact Gradient