发表机构
Fundamental AI Research (FAIR); Meta Superintelligence Labs(基础人工智能研究院(FAIR); Meta超级智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究同时使用一阶和二阶矩估计量的随机逼近方法,涵盖Adam和Muon等优化器,通过最优预条件视角推导方法,并利用两阶段框架证明其几乎必然收敛到目标解邻域,邻域大小取决于估计量的偏差和方差。
AI 中文摘要
经典随机逼近方法依赖于随机回归函数的一阶矩(均值)的估计量。我们研究同时使用一阶矩和二阶矩估计量的方法,这些方法将现代深度学习优化器(如Adam和Muon)作为特例包含在内。我们通过求解矩阵方程的最优预条件处理的视角推导二阶矩随机逼近方法,并为其收敛性分析开发了一个两阶段框架。第一阶段聚焦于依赖精确一阶矩和二阶矩的概念性(不切实际的)方法的分析。在第二阶段,我们用各自的估计量替换精确矩,并引用Dvoretzky定理来证明由此产生的实用方法几乎必然收敛到目标解的邻域。该邻域的大小取决于一阶矩和二阶矩估计量的偏差和方差。我们为Muon和Adam的一种谱变体推导了具体界限,这些界限决定了它们收敛邻域的半径。
英文摘要
Classical stochastic approximation methods rely on estimators of the first moment (mean) of a random regression function. We study methods that employ estimators of both the first and the second moments, which include modern deep-learning optimizers such as Adam and Muon as special cases. We derive second-moment stochastic approximation methods through the lens of optimal preconditioning for solving matrix equations, and develop a two-stage framework for their convergence analysis. The first stage focuses on the analysis of conceptual (impractical) methods that rely on the exact first and second moments. In the second stage, we replace the exact moments with their respective estimators, and invoke Dvoretzky's theorem to show that the resulting practical methods converge almost surely to a neighborhood of the target solution. The size of the neighborhood depends on the biases and variances of the first- and second-moment estimators. We derive concrete bounds for Muon and a spectral variant of Adam that determine the radius of their neighborhood of convergence.