镜像Polyak与原始-对偶提升
Mirror Polyak and a Primal-Dual Lifting
浏览论文内容
中文总结 AI 辅助
该研究提出镜像Polyak方法,将Polyak步长的自适应特性扩展到镜像下降法,还通过凸对偶的提升形式,使镜像Polyak在结构化优化问题中无需知晓最优值即可达到对应保证。
中文摘要 AI 辅助
一阶方法通常需要依赖目标函数正则性条件的特定步长,例如光滑性、Lipschitz连续性或强凸性常数。Polyak步长是凸函数次梯度下降的经典替代方案,仅利用目标函数最优值的信息,且能自动适配上述各类情况。然而,许多优化问题更适合用非欧几里得几何描述,且更易采用镜像下降法,将这种自适应特性扩展到镜像下降法中并非易事。现有部分Polyak步长的推广依赖于范数而非纯粹的相对几何,排除了镜像下降法的诸多应用场景。在本研究中,我们重新审视Kiwiel(1997)提出的基于Bregman投影的Polyak步长变体,将其命名为镜像Polyak。该方法已知能渐近收敛,但其收敛速率尚不明确。我们证明镜像Polyak享有与其欧几里得对应方法相似的保证,能自动适配相对意义下的光滑性、Lipschitz连续性或强凸性。随后,我们利用镜像Polyak避免在正则化线性回归和逻辑回归等结构化优化问题中需知晓最优值的要求。我们提出一种基于凸对偶的提升形式,其最优值恰好为零,且由问题结构给出自然的镜像映射。若已知最优值,将镜像Polyak应用于该提升问题时,可获得与原问题中Polyak步长相同的最坏情况保证。
英文摘要
First-order methods typically require a specific step-size that depends on the regularity conditions of the objective function, such as the smoothness, Lipschitz continuity, or strong convexity constants. The Polyak step-size is a classical alternative for subgradient descent on convex functions that only uses knowledge of the optimal value of the objective function and automatically adapts to the above-mentioned regimes. However, many optimization problems are better described by non-Euclidean geometries and are more amenable to mirror descent. Extending this adaptivity to mirror descent is subtle. Some existing generalizations of the Polyak step-size rely on norms instead of purely on relative geometry, excluding many of the use cases of mirror descent. In this work, we revisit a variant of the Polyak step-size based on Bregman projections due to Kiwiel (1997), which we call mirror Polyak. This method is known to converge asymptotically, but its convergence rate is not known. We show that mirror Polyak enjoys guarantees similar to its Euclidean counterpart, automatically adapting to relative notions of smoothness, Lipschitz continuity, or strong convexity. We then leverage mirror Polyak to avoid having to know the optimal value in some structured optimization problems such as regularized linear and logistic regression. We propose a lifted formulation based on convex duality with optimal value exactly zero and a natural mirror map given by the problem's structure. Mirror Polyak applied to the lifted problem enjoys the same worst-case guarantees as the Polyak step-size in the original problem if we knew the optimal value.