arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

随机梯度方法的最优两步步长调度

Optimal Two-Step Stepsize Schedule for Stochastic Gradient Methods

Luwei Bai, Baoyu Zhou

arXiv 2608.15035首次发表:更新:

AI 中文总结

该研究针对强凸光滑函数的随机梯度方法,刻画了依赖初始最优性间隙与噪声水平比值的全局最优两步步长调度,揭示了步长随噪声减弱而增大的平衡规律。

AI 中文摘要

结构化的非恒定大步长可提升确定性设置下梯度下降的收敛性,但在随机优化中,激进步长会放大 oracle 噪声,阻碍随机梯度方法的收敛。针对强凸且光滑的函数,仅假设能获取有限支撑、有界方差的无偏随机梯度估计,我们刻画了应用于该类函数的随机梯度方法的全局最优两步步长调度。该最优调度依赖初始最优性间隙与噪声水平的比值,呈现出多种不同的状态;随着随机噪声的影响减弱,最优两步步长会变大,体现了迭代收敛速度提升的益处与随机噪声引发的扰动之间的平衡。

英文摘要

Structured nonconstant large stepsizes can improve the convergence of gradient descent in the deterministic setting. However, in stochastic optimization, aggressive stepsizes can amplify oracle noise and hinder the convergence of stochastic gradient methods. We characterize the globally optimal two-step stepsize schedule for stochastic gradient methods applied to strongly convex and smooth functions, assuming access only to unbiased stochastic gradient estimates with finite support and bounded variance. The optimal schedule depends on the ratio of the initial optimality gap to the noise level and exhibits several distinct regimes. As the influence of stochastic noise diminishes, the optimal two-step stepsizes become larger, reflecting a balance between the benefits of faster iterate convergence and the perturbations induced by stochastic noise.

Comments23 pages, 2 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑