arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

平衡早期性能牺牲与长期收益:跨训练视野扩展学习率预热时长

Balancing Early Performance Sacrifices with Long-Term Gains: Scaling Learning-Rate Warmup Duration Across Training Horizons

Kristi Topollai, Anna Choromanska

arXiv 2609.33041首次发表:更新:

AI 中文总结

本研究提出一个二次模型,揭示学习率预热时长与训练视野及峰值学习率的关系,并推导出视野缩放定律,以预测长视野下的最优预热时长,建议将其作为依赖视野的超参数。

AI 中文摘要

学习率预热是语言模型训练中的一项标准技术,然而其持续时间在很大程度上仍依赖启发式方法。常见做法要么使用固定的更新次数,要么使用训练视野的固定比例,这两种选择在训练时间延长时意味着截然不同的扩展方式。预热何时应保持固定,何时应随视野增长?我们用一个二次模型来解决这个问题,该模型的不同模式对峰值学习率的响应各异。预热会减缓在峰值学习率下已经收缩良好的方向上的进展,但可以消除接近稳定性边缘方向上的持续误差,而更高的峰值学习率会将平衡转向更长的预热时长。这产生了一个紧凑的视野缩放定律,能够涵盖从基本无预热、固定时长预热到随训练视野增长的时长等不同区间,并解释了首选区间如何随峰值学习率变化。由于该定律捕捉了放弃早期进展与改善后续轨迹之间的权衡,因此可以使用较短的运行来拟合,并用于预测在显著更长视野下的预热情况。综合来看,我们的结果通过单一的权衡解释了预热的几个熟悉特性,并建议将预热时长视为一个依赖于视野的超参数,而非固定的训练启发式方法。

英文摘要

Learning-rate warmup is a standard technique in language-model training, yet its duration remains largely heuristic. Common approaches use either a fixed number of updates or a fixed fraction of the training horizon, two choices that imply very different scaling as training gets longer. When should warmup stay fixed, and when should it grow with the horizon? We address this question with a quadratic model whose modes respond differently to the peak learning rate. Warmup slows progress in directions that already contract well at the peak rate, but can remove persistent error in directions near the stability edge, with higher peak rates shifting the balance toward longer warmup durations. This yields a compact horizon scaling law that captures regimes ranging from essentially no warmup, through fixed-duration warmup, to durations that grow with the training horizon, and explains how the preferred regime changes with peak learning rate. Because the law captures the tradeoff between giving up early progress and improving the trajectory that follows, it can be fit using shorter runs and used to predict warmup at substantially longer horizons. Together, our results explain several familiar properties of warmup through a single tradeoff and suggest treating warmup duration as a horizon-dependent hyperparameter rather than a fixed training heuristic.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑