发表机构
Normal Computing; Cornell University; Stanford University(Normal Computing; 康奈尔大学; 斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究将幂律谱理论从确定性梯度下降推广至随机梯度下降,并针对各向同性高斯数据精确分析洗牌策略,证明Single Shuffle优于Flip-Flop和IID采样,弥合了谱理论与实际训练选择之间的差距。
AI 中文摘要
近期研究已确立,数据上的幂律谱条件能够为确定性梯度下降提供紧的收敛界,从而解决了经典指数界与观测到的幂律学习曲线之间的冲突。在本工作中,我们将此结果扩展到高维机器学习的随机机制。我们提供两项主要贡献:(1)我们将幂律谱理论推广到随机梯度下降(SGD),表明相同的谱指数控制随机动力学;(2)对于各向同性高斯数据这一基础情形,我们提供了数据洗牌的精确分析,推导出精确常数,证明Single Shuffle严格优于Flip-Flop和IID采样。我们的结果弥合了抽象谱理论与实际随机训练选择之间的差距,提供了数据几何如何驱动优化速度的统一图景。
英文摘要
Recent work has established that power-law spectral conditions on data enable tight convergence bounds for deterministic gradient descent, resolving the conflict between classical exponential bounds and observed power-law learning curves. In this work, we extend this result to the stochastic regime of high-dimensional machine learning. We provide two main contributions: (1) We generalize the power-law spectral theory to Stochastic Gradient Descent (SGD), showing that the same spectral exponents govern stochastic dynamics; (2) For the fundamental case of isotropic Gaussian data, we provide a precise analysis of data shuffling, deriving exact constants that prove Single Shuffle is strictly superior to Flip-Flop and IID sampling. Our results bridge the gap between abstract spectral theory and practical stochastic training choices, offering a unified picture of how data geometry drives optimization speed.
Comments59 pages, 9 figures, 2 tables