arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2405.14578cs.LG

最优学习率与批量大小缩放中的激增现象

Surge Phenomenon in Optimal Learning Rate and Batch Size Scaling

  • Tencent Hunyuan(腾讯混元)
  • Institute of Computational Social Science, Peking University(北京大学计算社会科学研究院)
  • University of Macau(澳门大学)

机构由 AI 辅助整理,请以论文原文为准。

Shuaipeng Li, Penghao Zhao, Hailin Zhang, Xingwu Sun, Hao Wu, Dian Jiao, Weiyan Wang, Chengjun Liu, Zheng Fang, Jinbao Xue, Yangyu Tao, Bin Cui, Di Wang

更新

AI总结:

本文针对Adam风格优化器,理论推导并实验验证最优学习率随批量大小先升后降的激增缩放定律及其峰值随训练向大批量移动的规律。

AI中文摘要:

在当前的深度学习任务中,Adam、Adagrad、RMSProp、Adafactor和Lion等Adam风格优化器已被广泛用作SGD风格优化器的替代方案。这些优化器通常使用梯度的符号来更新模型参数,从而产生更稳定的收敛曲线。学习率和批量大小是优化器最关键的超参数,需要仔细调优才能实现有效收敛。以往研究表明,对于SGD风格优化器,最优学习率随批量大小线性增加或遵循类似规律。然而,这一结论并不适用于Adam风格优化器。本文通过理论分析和大量实验,阐明了Adam风格优化器的最优学习率与批量大小之间的联系。首先,我们提出了梯度符号情形下批量大小与最优学习率之间的缩放定律,证明最优学习率随批量大小增大先上升后下降。此外,随着训练推进,该激增的峰值将逐渐向更大的批量大小移动。其次,我们在多种CV和NLP任务上进行实验,验证了该缩放定律的正确性。

英文摘要:

In current deep learning tasks, Adam style optimizers such as Adam, Adagrad, RMSProp, Adafactor, and Lion have been widely used as alternatives to SGD style optimizers. These optimizers typically update model parameters using the sign of gradients, resulting in more stable convergence curves. The learning rate and the batch size are the most critical hyperparameters for optimizers, which require careful tuning to enable effective convergence. Previous research has shown that the optimal learning rate increases linearly or follows similar rules with batch size for SGD style optimizers. However, this conclusion is not applicable to Adam style optimizers. In this paper, we elucidate the connection between optimal learning rates and batch sizes for Adam style optimizers through both theoretical analysis and extensive experiments. First, we raise the scaling law between batch sizes and optimal learning rates in the sign of gradient case, in which we prove that the optimal learning rate first rises and then falls as the batch size increases. Moreover, the peak value of the surge will gradually move toward the larger batch size as training progresses. Second, we conducted experiments on various CV and NLP tasks and verified the correctness of the scaling law.

↑