欧几里得球面上的高效优化:黎曼梯度对齐作为统一原理
Efficient Optimization on the Euclidean Sphere: Riemannian Gradient Alignment as a Unifying Principle
浏览论文内容
中文总结 AI 辅助
本文提出黎曼梯度对齐(RGA)作为统一结构条件,证明满足该条件时黎曼梯度下降可快速收敛,并验证了多种学习问题,确立了新问题类的可处理性。
中文摘要 AI 辅助
欧几里得球面约束优化问题或许代表了非凸约束问题中最基本的一类,对于这类问题,针对特定问题实例(如特征值问题和近期文献中出现的学习问题)存在高效的优化算法。然而,何时以及为何我们能够使用简单的一阶方法在球面上高效优化,目前尚不清楚。我们引入黎曼梯度对齐(RGA)作为实现快速收敛的统一结构条件。RGA是一种参数化的方向误差界,要求切向量与目标解向量充分对齐。我们证明,当RGA以阶数$r=1$成立时,采用递减步长的黎曼梯度下降线性收敛;对于$r>1$,收敛速率为$O(k^{-1/(2r-2)})$。该结果甚至适用于并非任何目标函数梯度的切向量场。我们的主要技术贡献是RGA的一族基于导数的充分条件。我们为一系列学习问题证明了这些条件,包括多实例学习/最大池化、高斯半空间、单指标模型和广义线性模型、相位恢复、完全正交字典学习以及(稀疏)PCA。所提供的示例涵盖了包括标准高斯、条件高斯以及更一般的结构化分布在内的分布族。所提供的保证涵盖了通常(测地线)非凸、可能非光滑且可能来自离散分布的目标函数。这些结果确立了先前未知误差界条件的问题类的可处理性,并统一、加强和推广了若干先前结果。
英文摘要
Euclidean sphere-constrained optimization problems represent perhaps the most basic class of nonconvex-constrained problems for which there exist efficient optimization algorithms for specific problem instances, such as the eigenvalue problems and learning problems arising in the recent literature. However, when and why we can efficiently optimize on the sphere using simple first-order methods is poorly understood. We introduce Riemannian Gradient Alignment (RGA) as a unifying structural condition enabling fast convergence. RGA is a parameterized directional error bound requiring a tangent vector to sufficiently align with a target solution vector. We prove that Riemannian Gradient Descent with diminishing step sizes converges linearly when RGA holds with order $r=1$, and at rate $O(k^{-1/(2r-2)})$ for $r>1$. The result applies even to tangent vector fields that are not gradients of any objective. Our main technical contribution is a family of derivative-based sufficient conditions for RGA. We prove these conditions for a range of learning problems, including multi-instance learning/max pooling, Gaussian halfspaces, single-index and generalized linear models, phase retrieval, complete orthogonal dictionary learning, and (sparse) PCA. The provided examples span distributional families that include standard Gaussian, conditionally Gaussian, and more general structured distributions. The provided guarantees cover objectives that are generally not (geodesically) convex, may be nonsmooth, and may arise from discrete distributions. These results establish tractability of problem classes where no prior error bound conditions were known, as well as unify, strengthen, and generalize several prior results.
发表机构
- UW-Madison(威斯康星大学麦迪逊分校)
- MIT(麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。