终端收缩平均揭示大语言模型预训练中的调度器-估计器交互
Terminal Shrinkage Averaging Reveals a Schedule-Estimator Interaction in LLM Pretraining
浏览论文内容
中文总结 AI 辅助
提出终端收缩平均(TSA)方法,在原始最终迭代与最近检查点平均间插值,分离调度器与估计器选择,在NanoChat实验中验证其提升验证质量并加速基准。
中文摘要 AI 辅助
大语言模型(LLM)预训练通常直接返回原始最终迭代。这耦合了两个设计选择:生成参数轨迹的学习率调度器,以及构建部署模型的估计器(例如原始最终迭代或检查点平均)。促进优化进展的调度器可能不同于最小化原始最终迭代变异的调度器。分离这些选择创造了一个机会,即在训练后期保持进展的同时减少返回模型的变异。为此,我们提出了终端收缩平均(TSA),它在原始最终迭代与最近检查点的平均之间进行插值,以平衡近期进展与终端变异。我们在局部二次近似下分析了TSA如何改变偏好的终端学习率调度器,并通过一系列受控的NanoChat实验测试了这种交互。最后,我们证明了所得收益可转移到深度为22的NanoChat中,其中组合的调度器和估计器提高了验证质量。一次合格的time-to-GPT-2运行也比我们实验中使用的公共基线完成得更快,提供了基准加速的初步证据。
英文摘要
Large language model (LLM) pretraining conventionally returns the raw final iterate. This couples two design choices: the learning-rate schedule that generates the parameter trajectory and the estimator that constructs the deployed model (e.g. the raw final iterate or a checkpoint average). A schedule that promotes optimization progress may differ from one that minimizes variation in the raw final iterate. Separating these choices creates an opportunity to maintain progress late in training while reducing variation in the returned model. To this end, we propose \emph{Terminal Shrinkage Averaging (TSA)}, which interpolates between the raw final iterate and the average of recent checkpoints to balance recent progress against terminal variation. We analyze how TSA changes the preferred terminal learning-rate schedule under a local quadratic approximation and test this interaction through a sequence of controlled NanoChat experiments. Finally, we demonstrate that the resulting gains transfer to depth-22 NanoChat, where the combined schedule and estimator improve validation quality. A qualifying time-to-GPT-2 run also finishes faster than the public baseline used in our experiments, providing preliminary evidence of benchmark acceleration.
发表机构
- University of Michigan(密歇根大学)
机构由 AI 辅助整理,请以论文原文为准。