超越单位激励与光滑性的随机鞍点规避:一种路径式李雅普诺夫-佩龙框架
Stochastic Saddle Avoidance Beyond Unit Excitation and Smoothness: A Pathwise Lyapunov-Perron Framework
- National University of Singapore(新加坡国立大学)
- The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出路径式李雅普诺夫-佩龙框架,在不满足单位激励假设的常见随机优化场景下,证明随机递推的严格鞍点规避结果,确立相关算法收敛到原目标函数的局部极小值点。
AI中文摘要:
单位激励(UE)是随机鞍点规避中的常见假设:随机误差在期望意义下沿每个方向必须有一个一致正的分量。该条件为排除收敛到严格鞍点提供了直接方法,但它也过度简化了实际噪声结构,与许多随机优化场景不匹配。在过参数化或插值模型中,噪声可能在平稳性附近消失;在有限和问题中,随机梯度噪声可能位于低维、依赖数据的子空间中,这些常见场景下UE自然不成立。本文针对不满足UE的随机递推,证明了一个抽象的几乎必然规避定理,该定理用可验证的路径式条件替代UE型要求。在应用中,这些条件可由标准独立同分布采样下的局部光滑性和有限矩假设,或无放回采样下的有限和结构得出。由于随机采样的映射通常不共享不动点,确定性分析中著名的中心稳定流形论证无法直接应用,取而代之的是使用路径依赖的变量变换结合基于路径式李雅普诺夫-佩龙的证明策略。作为应用,我们得到了随机镜像下降(包括SGD)和随机重排的严格鞍点规避结果;对于非光滑复合目标,我们证明了一类近端型随机梯度方法的规避结果。将这些见解与合适的迭代收敛保证结合,可确立收敛到原目标函数的局部极小值点。
英文摘要:
Unit excitation (UE) is a common assumption in stochastic saddle avoidance: the stochastic error must have a uniformly positive component along every direction, in expectation. This condition gives a direct way to rule out convergence to strict saddles, but it also oversimplifies the actual noise structure, and does not match many stochastic optimization regimes. In overparameterized or interpolation models, the noise may vanish near stationarity. In finite-sum problems, the stochastic gradient noise may lie in a low-dimensional, data-dependent subspace. In these (common) scenarios, UE is naturally not satisfied. In this paper, we prove an abstract almost sure avoidance theorem for stochastic recursions without UE. The theorem replaces UE-type requirements by verifiable pathwise conditions. In applications, these conditions follow, e.g., from local smoothness and finite-moment assumptions under standard i.i.d. sampling, or from the finite-sum structure under without-replacement sampling. Since the stochastically sampled maps generally do not share a fixed point, the celebrated center-stable manifold argument used in deterministic analyses is not directly applicable. Instead, we use a path-dependent change of variables together with a pathwise Lyapunov--Perron-based proof strategy. As applications, we obtain strict saddle avoidance for stochastic mirror descent (including SGD) and for random reshuffling. For nonsmooth composite objectives, we prove avoidance results for a proximal-type stochastic gradient method. Combining these insights with suitable iterate convergence guarantees, this allows establishing convergence to local minimizers of the original objective function.