arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

神经网络优化器背后的原理

On the Principles Behind Neural Network Optimizers

Yushun Zhang

arXiv 2608.16760首次发表:更新:

发表机构

School of Data Science, The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)数据科学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本论文为优化器Adam建立了有原则的基础,通过海森矩阵结构分析其在Transformer上优于SGD的原因,提出内存减半且性能不变的新优化器Adam-mini,还为其他NN优化器提供了新见解。

AI 中文摘要

可靠的优化是神经网络(NN)训练的核心,但现代大语言模型(LLM)的默认优化器Adam却建立在脆弱的基础之上。本论文为Adam开发了有原则的基础并推动新设计。首先,我们重新审视Adam的发散-收敛争论,证明存在依赖于问题的相变:通过选择合适的、与批次大小相关的超参数,Adam会收敛;而在小β₂ regime下它可能发散。其次,我们通过海森矩阵(Hessian)结构研究Adam在Transformer上显著优于随机梯度下降(SGD)的原因,发现海森矩阵在训练过程中演变为近块对角形式,伴随强烈的块异质性,我们证明这种结构使Adam的对角预条件器有效,还进一步表明这种特殊海森矩阵结构源于大型矩阵变量的连续乘法,并基于随机矩阵理论提供了严格分析。最后,这些见解催生了新优化器Adam-mini,它在保持Adam性能的同时将内存占用减少50%。我们的结果还有超出Adam的更广泛意义:它们揭示了基于矩阵的非凸问题中的新局部结构,也有助于理解和改进近期的神经网络优化器,如Muon。

英文摘要

Reliable optimization is central to neural network (NN) training, yet Adam, the default optimizer for modern LLMs, rests on a fragile foundation. This thesis develops a principled grounding for Adam and motivates new designs. First, we revisit Adam's divergence--convergence debate and show the existence of a problem-dependent phase transition: with properly chosen, batch-size-dependent hyperparameters, Adam converges, whereas under small-$β_2$ regimes it can diverge. Second, we investigate why Adam substantially outperforms SGD on Transformers through Hessian structure. We find that the Hessian evolves toward a near-block-diagonal form along training, accompanied by strong block heterogeneity. We prove that this structure makes Adam's diagonal preconditioner effective. We further show that this special Hessian structure originates from consecutive multiplications of large matrix variables, and we provide a rigorous analysis based on random matrix theory. Finally, these insights motivate Adam-mini, a new optimizer that reduces Adam's memory footprint by 50\% while preserving its performance. Our results also have broader implications beyond Adam: they reveal new local structures in matrix-based nonconvex problems, and also help understand and improve recent NN optimizers, such as Muon.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑