arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

局部化、重启、加速:广义光滑性下的随机优化

Localize, Restart, Accelerate: Stochastic Optimization under Generalized Smoothness

Darina Dvinskikh, Alexander Gasnikov, Aleksandr Lobanov, Ilgam Latypov

arXiv 2609.06555首次发表:更新:

AI 中文总结

针对非对称广义光滑性下的随机凸优化,提出两阶段加速方法ARC-SG,通过局部化与重启认证实现加速收敛并恢复经典速率。

AI 中文摘要

我们研究了在非对称\((L_0,L_1)\)-广义光滑性下的随机凸优化,该模型受机器学习目标的启发,其局部曲率可能随梯度范数增长。我们假设一个具有加性范数-次高斯噪声的无偏一阶预言机。在此设置中加速是困难的,因为动量可能进入曲率大得多的区域,而随机梯度无法可靠地认证无限制的轨迹。我们提出了\ extsf{ARC-SG},一种两阶段加速方法:第一阶段使用广义光滑性感知的随机步骤减少过大的梯度,第二阶段通过限制在认证光滑性球内的重启加速求解器来求解强凸近端子问题。精确近端点不会增加梯度范数,使得这些认证能够通过外层循环传播。贡献在于一种逐查询认证局部化构造,具有显式的广义光滑性因子和强凸重启扩展。\ extsf{ARC-SG}以高概率实现加速优化贡献和光滑子类最优的精度统计依赖性,直至对数和广义光滑性因子。其凸精度指数与在更广泛光滑性和仿射方差模型下同时期的公开随机加速结果一致;我们的区别在于认证几何、显式参数核算和强凸保证。当\(L_1=0\)时,结果恢复经典的加速随机速率。在具有无界梯度的目标上的实验展示了两阶段机制及其有限预算优势。

英文摘要

We study stochastic convex optimization under asymmetric \((L_0,L_1)\)-generalized smoothness, a model motivated by machine-learning objectives whose local curvature may grow with the gradient norm. We assume an unbiased first-order oracle with additive norm-sub-Gaussian noise. Acceleration is difficult in this setting because momentum may enter regions of much larger curvature, while stochastic gradients cannot reliably certify an unrestricted trajectory. We propose \textsf{ARC-SG}, a two-phase accelerated method: Phase~I reduces excessively large gradients using a generalized-smoothness-aware stochastic step, then Phase~II solves strongly convex proximal subproblems by a restarted accelerated solver confined to certified smoothness balls. Exact proximal points do not increase the gradient norm, allowing these certificates to propagate through the outer loop. The contribution is a query-by-query certified-localization construction with explicit generalized-smoothness factors and a strongly convex restart extension. \textsf{ARC-SG} achieves, with high probability, an accelerated optimization contribution and smooth-subclass-optimal statistical dependence on accuracy, up to logarithmic and generalized-smoothness factors. Its convex accuracy exponents agree with a contemporaneous public stochastic-acceleration result under a broader smoothness and affine-variance model; our distinction is the certified geometry, explicit parameter accounting, and strongly convex guarantee. The results recover classical accelerated stochastic rates when \(L_1=0\). Experiments on objectives with unbounded gradients illustrate the two-phase mechanism and its finite-budget advantage.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑