arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大批量训练中的动量:Polyak增大临界批量规模,Nesterov提升数据效率

Momentum in large-batch training: Polyak enlarges the critical batch size, Nesterov improves data efficiency

Jia-Nan Wang, Zixun Huang, Kairui Li, Lei Wu

arXiv 2609.02728首次发表:更新:

发表机构

Peking University; University of Pennsylvania; AI for Science Institute, Beijing(北京大学; 宾夕法尼亚大学; 北京科学智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究以幂律核回归为场景,分析单轮训练中Polyak和Nesterov动量对大批量训练的影响,推导临界学习率与风险动态标度律,得出三区域批量相图,验证了Polyak扩临界批量、Nesterov提数据效率的结论。

AI 中文摘要

我们以幂律核回归为易处理的研究场景,探究在单轮训练模式下动量何时以及如何改善大批量训练效果。我们首先通过临界学习率(定义为稳定训练的最大学习率)刻画风险稳定性,得到$η_{\text{SGD}}^{\text{crit}}\eqsim 1$、$η_{\text{Polyak}}^{\text{crit}}\eqsim \min\{1,B(1-ρ)\}$、$η_{\text{Nesterov}}^{\text{crit}}\eqsim \min\{1,B^β(1-ρ)\}$,其中$B$为批量规模,$ρ$为动量因子,$β>1$为容量指数。在该可行区域内,我们推导了完整风险动态的标度律,捕捉了从早期瞬态、经幂律衰减到噪声基底的演化过程。随后我们在固定数据预算下,在可行学习率和动量因子范围内最小化最终步风险,得到三区域批量规模相图,揭示了动量的作用如何随批量规模变化。值得注意的是,Polyak增大了临界批量规模——即能保持最优小批量数据标度指数的最大批量规模,从而在不牺牲数据效率的前提下实现更高并行度。相比之下,Nesterov在大批量场景下实现了更优的数据效率,因为其前瞻机制抑制了噪声累积。数值实验验证了预测的稳定性边界、风险动态和批量规模相图。

英文摘要

We study when and how momentum improves large-batch training in the one-pass regime, using power-law kernel regression as a tractable setting. We first characterize risk stability through the critical learning rate, defined as the largest learning rate for stable training, and obtain $η_{\mathrm{SGD}}^{\mathrm{crit}}\eqsim 1$, $η_{\mathrm{Polyak}}^{\mathrm{crit}}\eqsim \min\{1,B(1-ρ)\}$, and $η_{\mathrm{Nesterov}}^{\mathrm{crit}}\eqsim \min\{1,B^β(1-ρ)\}$, where $B$ is the batch size, $ρ$ is the momentum factor, and $β>1$ is the capacity exponent. Within this admissible region, we derive scaling laws for the full risk dynamics, capturing the progression from an early transient, through power-law decay, to a noise floor. We then minimize the final-step risk over the admissible learning rates and momentum factors under a fixed data budget, yielding a three-regime batch-size phase diagram that reveals how the role of momentum changes with batch size. Notably, Polyak enlarges the critical batch size, the largest batch size preserving the best small-batch data-scaling exponent, thereby enabling greater parallelism without sacrificing data efficiency. In contrast, Nesterov achieves better data efficiency in the large-batch regime because its look-ahead mechanism suppresses noise accumulation. Numerical experiments validate the predicted stability boundaries, risk dynamics, and batch-size phase diagram.

Comments69 pages, 8 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑