发表机构
Tsinghua University; Peking University; Shanghai Jiao Tong University(清华大学; 北京大学; 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文从局部景观几何角度分析语言模型预训练,揭示两阶段动态,解释学习率预热和动态批次大小调度,提供大规模预训练调优策略。
AI 中文摘要
预训练语言模型的规模和成本使得高效的超参数调优至关重要,然而目前仍缺乏原则性的指导。在本工作中,我们从局部景观几何的角度分析语言模型预训练动态。我们的研究揭示了两个不同的阶段。在第一阶段,局部景观的锐度最初较高,导致在大学习率下出现不稳定和损失平台期。在训练早期,景观从锐利区域向平坦区域转变。这一动态解释了学习率预热的必要性,并进一步表明较大的峰值学习率需要相应更长的预热期。在第二阶段,局部景观由梯度噪声尺度主导。我们的理论识别出一个深度平坦度权衡:较小批次带来的高噪声拓宽了损失盆地,而较大批次降低噪声则加深了损失盆地。该理论启发了动态批次大小调度器,其从小批次开始,并在训练后期增大批次大小。综合起来,我们提供了损失景观演化的统一视角,转化为大规模预训练的可操作调优策略。
英文摘要
The scale and expense of pre-training language models make efficient hyperparameter tuning essential, yet a principled guidance is still missing. In this work, we analyze language model pre-training dynamics from a local landscape geometry perspective. Our study reveals two distinct phases. In Phase I, sharpness of the local landscape is initially high, leading to instability and loss plateaus under large learning rates (LRs). The landscape shifts from sharp to flatter regions early in training. This dynamic explains the necessity of LR warmup and further suggests that larger peak LRs require proportionally longer warmup periods. In Phase II, the local landscape is governed by the gradient noise scale. Our theory identifies a depth flatness trade-off: high noise from smaller batches widens the loss basin, whereas reduced noise from larger batches deepens it. This theory motivates a dynamic batch-size (BS) scheduler that begins with a small BS and increases it late in training. Together, we provide a unified view of loss landscape evolution, which translates into actionable tuning strategies for large-scale pre-training.
Comments23 pages, 15 figures