AI 中文总结
本文提出随机在线缩放梯度方法(SOSGM),推广自适应预条件框架至随机优化,兼容多种预条件器与动量,成本与Adam相当,实验表明其性能优于现有自适应一阶方法。
AI 中文摘要
本文提出了随机在线缩放梯度方法(SOSGM),这是对最近在arXiv:2505.23081和arXiv:2509.11007中提出的自适应预条件框架在随机优化中的推广。在标准假设下,我们利用大批量或方差缩减技术建立了SOSGM的收敛性保证。SOSGM兼容流行的对角和/或低秩预条件器以及重球动量,同时保持与Adam相当的内存和计算成本。广泛的数值实验证明了SOSGM强大的实证性能。使用对角预条件器时,SOSGM及其变体在一系列统计学习任务中显著优于现有的自适应一阶方法。
英文摘要
This paper introduces Stochastic Online Scaled Gradient Methods (SOSGM), a generalization of the recently developed adaptive preconditioning framework in arXiv:2505.23081 and arXiv:2509.11007 to stochastic optimization. Under standard assumptions, we establish convergence guarantees for SOSGM using large batchsize or variance reduction. SOSGM is compatible with popular diagonal and/or low-rank preconditioners as well as heavy-ball momentum, while maintaining memory and computation cost comparable to Adam. Extensive numerical experiments demonstrate the strong empirical performance of SOSGM. Using a diagonal preconditioner, SOSGM and its variants substantially outperform existing adaptive first-order methods across a range of statistical learning tasks.