arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2407.07972cs.LGcs.AI

解构什么造就了语言模型的好优化器

Deconstructing What Makes a Good Optimizer for Language Models

  • Harvard University(哈佛大学)
  • Kempner Institute at Harvard University(哈佛大学肯普纳研究所)

机构由 AI 辅助整理,请以论文原文为准。

Rosie Zhao, Depen Morwani, David Brandfonbrener, Nikhil Vyas, Sham Kakade

更新

AI总结:

本研究比较多种优化器在语言模型中的表现,发现除SGD外性能相当,并剖析Adam简化版本,揭示最后一层和LayerNorm的适应性对性能与稳定性至关重要。

AI中文摘要:

训练语言模型的成本随着规模的扩大而日益增加,这促使人们多次尝试提高优化效率。尽管有这些努力,Adam优化器仍然是最广泛使用的,因为普遍认为它是最有效的方法。我们旨在比较几种优化算法,包括SGD、Adafactor、Adam、Lion和Sophia,在自回归语言建模的背景下,跨越一系列模型规模、超参数和架构变体。我们的研究结果表明,除了SGD之外,这些算法在最优性能和广泛超参数选择下的表现方面都相当。我们的结果向从业者建议,优化器的选择可以由实际考虑因素(如内存约束和实现简便性)来指导,因为没有单一算法在性能或对超参数错误指定的稳定性方面成为明显赢家。鉴于我们的发现,我们进一步剖析这些方法,检查Adam的两种简化版本:a) 符号动量(Signum),我们发现它恢复了Adam的性能和超参数稳定性;b) Adalayer,我们引入的Adam的逐层变体,以研究Adam对网络不同层的预处理的影响。检查Adalayer使我们得出结论,也许令人惊讶的是,特别是在最后一层和LayerNorm参数上的适应性对于保持性能和学习率稳定性是必要的。

英文摘要:

Training language models becomes increasingly expensive with scale, prompting numerous attempts to improve optimization efficiency. Despite these efforts, the Adam optimizer remains the most widely used, due to a prevailing view that it is the most effective approach. We aim to compare several optimization algorithms, including SGD, Adafactor, Adam, Lion, and Sophia in the context of autoregressive language modeling across a range of model sizes, hyperparameters, and architecture variants. Our findings indicate that, except for SGD, these algorithms all perform comparably both in their optimal performance and also in terms of how they fare across a wide range of hyperparameter choices. Our results suggest to practitioners that the choice of optimizer can be guided by practical considerations like memory constraints and ease of implementation, as no single algorithm emerged as a clear winner in terms of performance or stability to hyperparameter misspecification. Given our findings, we further dissect these approaches, examining two simplified versions of Adam: a) signed momentum (Signum) which we see recovers both the performance and hyperparameter stability of Adam and b) Adalayer, a layerwise variant of Adam which we introduce to study the impact on Adam's preconditioning for different layers of the network. Examining Adalayer leads us to the conclusion that, perhaps surprisingly, adaptivity on both the last layer and LayerNorm parameters in particular are necessary for retaining performance and stability to learning rate.

补充信息

↑