arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

非凸非凹极小极大景观的黎曼上升-下降:盆地鞍点的收敛性及其在分布鲁棒优化中的应用

Riemannian ascent--descent for nonconvex nonconcave minimax landscapes: convergence to basin saddle points and applications to distributionally robust optimization

Rishabh Dixit, Pranav Upadrashta, Alex Cloninger

arXiv 2609.14141首次发表:更新:

发表机构

UC San Diego(加州大学圣地亚哥分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对非凸非凹极小极大景观,提出盆地鞍点概念及黎曼梯度上升多步下降迭代的收敛框架,应用于高斯分布鲁棒优化,获得线性或多项式收敛速率。

AI 中文摘要

我们研究了一类针对统计风险问题的分布鲁棒优化(DRO)问题,其被表述为欧几里得空间与黎曼流形乘积上的极小极大问题。由于所得极小极大景观通常是非凸非凹的,目前尚无已知的全局收敛的一阶方法。我们转而引入了“盆地鞍点”的概念,这是一种在局部极小值临界集的连通分量周围的δ盆地与测度流形上的测地球的笛卡尔积上局部定义的纳什均衡。我们在局部极小值临界集的连通分量周围的δ盆地中,在指数β∈(1,2]的局部Łojasiewicz型增长条件下,为黎曼梯度上升多步下降迭代发展了一个抽象收敛框架,使其收敛到盆地鞍点。在临界集的Lipschitz正则性下,我们建立了β=2时的线性收敛和β∈(1,2)时的多项式收敛到盆地鞍点,并明确依赖于流形的截面曲率。然后,我们将该框架应用于高斯测度上的统计风险DRO问题,其中模糊集自然建模为欧几里得空间与协方差矩阵的Bures Wasserstein流形的乘积,我们将其松弛为惩罚DRO公式。我们推导了所得拉格朗日函数的非渐近Hessian估计,建立了其极大值的存在性和局部唯一性,并证明了交替黎曼梯度方案收敛到惩罚DRO问题的盆地鞍点,恢复了抽象理论的线性和多项式速率,所有常数均以数据维度、损失矩和参考协方差显式表示。

英文摘要

We study a class of distributionally robust optimization (DRO) problems for the statistical risk problem, formulated as minimax problems over the product of a Euclidean space and a Riemannian manifold. Because the resulting minimax landscape is nonconvex nonconcave in general, no globally convergent first order method is known to be available. We instead introduce the notion of a \emph{basin saddle point}, a Nash equilibrium defined locally on the Cartesian product of a $δ$ basin around a connected component of the local minima critical set and a geodesic ball on the measure manifold. We develop an abstract convergence framework for a Riemannian gradient ascent multistep descent iteration to a basin saddle point under a local Łojasiewicz type growth condition, with exponent $β\in (1,2]$, in the $δ$ basin around connected components of the local minima critical sets. Under Lipschitz regularity of critical sets we establish linear convergence for $β= 2$ and polynomial convergence for $β\in (1,2)$ to a basin saddle point, with explicit dependence on the sectional curvature of the manifold. We then instantiate this framework for the statistical risk DRO problem over Gaussian measures, where the ambiguity set is naturally modeled as the product of Euclidean space and the Bures Wasserstein manifold of covariance matrices, which we relax to a penalized DRO formulation. We derive nonasymptotic Hessian estimates for the resulting Lagrangian, establish existence and local uniqueness of its maximizer, and prove that an alternating Riemannian gradient scheme converges to a basin saddle point of the penalized DRO problem, recovering the linear and polynomial rates of the abstract theory with all constants explicit in terms of data dimension, loss moments, and the reference covariance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑