arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

幂律谱下的随机优化:紧界与洗牌分析

Stochastic Optimization Under Power-Law Spectra: Tight Bounds and Shuffling Analysis

Thomas Dybdahl Ahle, Yaroslav Bulatov, Christopher De Sa, Christopher Ré

arXiv 2609.36271首次发表:更新:

发表机构

Normal Computing; Cornell University; Stanford University(Normal Computing; 康奈尔大学; 斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究将幂律谱理论从确定性梯度下降推广至随机梯度下降,并针对各向同性高斯数据精确分析洗牌策略,证明Single Shuffle优于Flip-Flop和IID采样,弥合了谱理论与实际训练选择之间的差距。

AI 中文摘要

近期研究已确立,数据上的幂律谱条件能够为确定性梯度下降提供紧的收敛界,从而解决了经典指数界与观测到的幂律学习曲线之间的冲突。在本工作中,我们将此结果扩展到高维机器学习的随机机制。我们提供两项主要贡献:(1)我们将幂律谱理论推广到随机梯度下降(SGD),表明相同的谱指数控制随机动力学;(2)对于各向同性高斯数据这一基础情形,我们提供了数据洗牌的精确分析,推导出精确常数,证明Single Shuffle严格优于Flip-Flop和IID采样。我们的结果弥合了抽象谱理论与实际随机训练选择之间的差距,提供了数据几何如何驱动优化速度的统一图景。

英文摘要

Recent work has established that power-law spectral conditions on data enable tight convergence bounds for deterministic gradient descent, resolving the conflict between classical exponential bounds and observed power-law learning curves. In this work, we extend this result to the stochastic regime of high-dimensional machine learning. We provide two main contributions: (1) We generalize the power-law spectral theory to Stochastic Gradient Descent (SGD), showing that the same spectral exponents govern stochastic dynamics; (2) For the fundamental case of isotropic Gaussian data, we provide a precise analysis of data shuffling, deriving exact constants that prove Single Shuffle is strictly superior to Flip-Flop and IID sampling. Our results bridge the gap between abstract spectral theory and practical stochastic training choices, offering a unified picture of how data geometry drives optimization speed.

Comments59 pages, 9 figures, 2 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑