arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30957cs.LGmath.OC

理解动量在河谷损失景观中的加速作用

Towards Understanding Momentum Acceleration in River-Valley Loss Landscape

Miao Lu, Zeyu Bian, Kaiyue Wen, Beining Wu, Siyu Chen, Tianhao Wang, Zhiyuan Li

首次发表
浏览论文内容

中文总结 AI 辅助

本研究在河谷损失景观下理论分析动量加速机制,发现动量通过稳定大学习率提升沿河优化速度,且对平坦缓转河谷,加速主要源于更大学习率而非动量本身。

中文摘要 AI 辅助

预训练大型语言模型的经验成功激发了对底层损失景观和优化动态的更深入探究。近期的实证和理论研究提出,训练损失景观常呈现“河谷”结构,其特征是存在一个低损失流形(河),两侧是损失较高的陡峭正交方向(山)。长期来看,优化进展主要由沿河的进展决定。在这样的景观中,使用大学习率的梯度下降可以沿河更快移动,尽管由于垂直振荡导致表观损失较高;而随后学习率的急剧衰减抑制了这些振荡,揭示了真正的优化进展。这解释了近期warmup-stable-decay(WSD)学习率调度器的成功,与余弦调度不同,它保持稳定的高学习率并在产生中间检查点之前衰减。在此基础之上,本工作进一步研究动量在此类损失景观中的作用。我们建立了理论分析,刻画动量如何通过稳定大学习率来加速优化,而大学习率是普通梯度下降无法容忍且不会显著偏离河的。由此启用的大学习率反过来沿河提供更大速度,并在长期内实现更快的本质进展。理论中另一个有趣的观察是,对于具有非常平坦且缓慢旋转的河的河谷景观,动量本身并不直接贡献于沿河跟踪速度的加速,而主要加速来自可接受的更大学习率。

英文摘要

The empirical success of pretraining large language models has inspired a deeper investigation into the underlying loss landscapes and the optimization dynamics. Recent empirical and theoretical study suggest that the training loss landscape often exhibits a "river-valley" structure, which features a low-loss manifold (river) flanked by sharp orthogonal directions with higher loss (mountains). In the long term, the optimization progress is determined primarily by the progress along the river. Within such a landscape, gradient descent with large learning rates can move faster along the river despite high apparent loss due to vertical oscillations, while a subsequent sharp decay in the learning rate suppresses these oscillations, revealing genuine optimization progress. This explains the recent success of warmup-stable-decay (WSD) learning rate scheduler which, unlike cosine scheduling, keeps stable high learning rate and decays before producing intermediate checkpoints. Building on this foundation, in this work we take a step further and study the role of momentum within such a loss landscape. We establish theoretical analysis that characterizes how momentum accelerates optimization by stabilizing large learning rates that can not be tolerated by vanilla GD without deviating significantly from the river. The enabled large learning rate in-turn gives greater speed along the river and makes faster essential progress in the long run. Another intriguing observation from theory is that for a river-valley landscape with very flat and slow-spinning river, the momentum itself does not contribute directly to acceleration in terms of the speed of tracking the river, while the main acceleration comes from the admissible larger learning rate.

发表机构

  • Stanford University(斯坦福大学)
  • University of California, San Diego(加州大学圣迭戈分校)
  • University of Chicago(芝加哥大学)
  • Yale University(耶鲁大学)
  • Toyota Technological Institute at Chicago(芝加哥丰田技术学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑