arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25551cs.LGmath.OCmath.STstat.MLstat.TH

随机优化中超越最优速率:轨迹自适应停止规则

Beyond Optimal Rates in Stochastic Optimization: Trajectory-Adaptive Stopping Rules

Liviu Aolaritei, Lucas Lévy, Francis Bach, Michael I. Jordan

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对强凸随机优化,提出轨迹自适应停止规则,构建相关置信序列与不等式,将其扩展至小批量SGD,实验显示该规则所需迭代次数远少于确定时间范围。

中文摘要 AI 辅助

随机梯度下降(SGD)通常在算法运行前选定确定的时间范围进行分析,而实际的停止决策是通过检查不断演化的轨迹自适应做出的。这种不匹配产生了一个根本性的验证问题:固定时间的保证通常在依赖数据的停止时间不再有效,而由最坏情况边界导出的确定时间范围可能极具保守性。我们针对强凸随机优化解决该问题,通过构造完全可观测、轨迹自适应的上置信序列,用于最后一次迭代到最优解的平方距离以及加权平均的次优性。这些边界随时间同时成立,在最坏情况下达到最优的1/t衰减速率(迭代对数因子以内),并适应已实现的随机梯度,允许SGD在验证规定精度后立即停止,同时不牺牲统计有效性。我们的方法将演化的SGD轨迹视为序贯实验,其观测结果提供关于未知优化误差的证据。为形式化这一视角,我们开发了新的递归置信序列技术和适用于随时间变化的条件均值及可预测范围(可能无界增长)的自适应过程的通用时间一致经验伯恩斯坦不等式。我们进一步将这些置信序列构造扩展到小批量SGD,其中经验伯恩斯坦边界利用每个小批量内已实现的二阶矩结构。数值实验表明,所得停止规则所需的迭代次数比自然确定时间范围少几个数量级。

英文摘要

Stochastic gradient descent (SGD) is typically analyzed at a deterministic horizon chosen before the algorithm is run, even though practical stopping decisions are made adaptively by inspecting the evolving trajectory. This mismatch creates a fundamental certification problem: fixed-time guarantees do not generally remain valid at data-dependent stopping times, while deterministic horizons derived from worst-case bounds can be highly conservative. We address this problem for strongly convex stochastic optimization by constructing fully observable, trajectory-adaptive upper confidence sequences for the squared distance of the last iterate to the optimizer and the suboptimality of a weighted average. These bounds hold simultaneously over time, attain the optimal $1/t$ decay rate up to iterated-logarithmic factors in the worst case, and adapt to the realized stochastic gradients, allowing SGD to stop as soon as a prescribed accuracy is certified without sacrificing statistical validity. Our approach treats the evolving SGD trajectory as a sequential experiment whose observations provide evidence about the unknown optimization error. To formalize this perspective, we develop new recursive confidence-sequence techniques and a general time-uniform empirical Bernstein inequality for adapted processes with time-varying conditional means and predictable ranges that may grow without bound. We further extend these confidence-sequence constructions to minibatch SGD, with the empirical Bernstein bounds exploiting the realized second-moment structure within each minibatch. Numerical experiments show that the resulting stopping rules can require several orders of magnitude fewer iterations than natural deterministic horizons.

发表机构

  • UC Berkeley(加州大学伯克利分校)
  • Inria(法国国家信息与自动化研究所)
  • École Normale Supérieure(巴黎高等师范学院)
  • PSL Research University(巴黎文理研究大学)
  • Department of Statistics, UC Berkeley(加州大学伯克利分校统计系)

机构由 AI 辅助整理,请以论文原文为准。

↑