发表机构
University of Pennsylvania(宾夕法尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究提出WSqD学习率调度方法,受随机凸优化启发,用移位反平方根基础取代WSD恒定稳定阶段,保留最终线性冷却,其基础学习率调度与训练轮次无关且收敛速率最优,实验显示在语言模型预训练中表现良好。
AI 中文摘要
标准学习率调度方法(如余弦退火)与固定训练轮次相关,限制了其适应事后延长训练轮次的能力。热身-稳定-衰减(WSD)部分解决了这个问题,但它的峰值学习率仍基于原始训练轮次调整,训练延长时可能次优。受随机凸优化启发,我们提出WSqD(带平方根基础和线性衰减的热身),用移位的反平方根基础取代WSD的恒定稳定阶段,保留最终线性冷却。在随机凸设置下,WSqD可证明达到极小极大最优的最后迭代收敛速率$O(1/\sqrt{T})$。重要的是,其基础学习率调度与训练轮次无关,仅需训练轮次确定何时开始最终冷却。实验表明,在使用SlimPajama语料库进行语言模型预训练时,WSqD在多个训练轮次上匹配或优于精心调整的WSD和其他基线,同时复用单个峰值学习率。
英文摘要
Standard learning rate schedules such as cosine annealing are tied to a fixed training horizon, limiting their ability to accommodate post hoc horizon extension. Warmup-stable-decay (WSD) partially addresses this issue by maintaining a long constant-rate phase before a short linear cooldown, allowing training to resume from a pre-decay checkpoint. However, its peak learning rate is still tuned based on the original training horizon and can become suboptimal when training is extended. Motivated by stochastic convex optimization, we propose WSqD (Warmup with Square-root base and linear Decay), a learning rate schedule that replaces WSD's constant stable phase with a shifted inverse-square-root base while retaining the final linear cooldown. In the stochastic convex setting, WSqD provably attains the minimax-optimal $O(1/\sqrt{T})$ last-iterate convergence rate. Importantly, its base learning rate schedule is horizon-independent, and the training horizon is needed only to determine when to begin the final cooldown. Empirically, on language-model pretraining using the SlimPajama corpus, WSqD matches or outperforms carefully tuned WSD and other baselines across multiple training horizons while reusing a single peak learning rate.