有效学习率支配语言模型预训练中的损失动态
Effective Learning Rate Governs Loss Dynamics in Language Model Pretraining
浏览论文内容
中文总结 AI 辅助
该研究发现语言模型预训练中存在ELR崩溃现象,确定ELR是连接LR调度、范数控制与损失动态的共同坐标,提出的基于ELR的FSL可实现跨范数控制方法的迁移并解释延迟加速效应。
中文摘要 AI 辅助
我们在语言模型预训练中发现了有效学习率(ELR)崩溃现象:学习率(LR)与参数范数主要通过二者的比值即有效学习率(ELR)来支配损失动态。当不同训练运行的ELR匹配时,尽管LR和参数范数存在显著差异,整个训练过程的损失轨迹仍会发生崩溃。在不同优化器、架构、数据集和模型规模下,平均崩溃误差通常为10^-3量级,低于代表性配置中种子到种子的变异水平。系统消融实验确定了归一化设计以及LR-范数变化的时间尺度是崩溃精度的关键决定因素。受控干预进一步表明,权重衰减和Hyperball主要通过其诱导的ELR调度来影响损失动态。用ELR替代LR可使拟合的函数缩放律(FSL)在范数控制方法间迁移,所得基于ELR的FSL还能解释延迟加速这一范数控制的反复出现的效应。综上,这些结果确立ELR为连接LR调度、范数控制与损失动态的共同坐标。
英文摘要
We uncover ELR collapse in language model pretraining: learning rate (LR) and parameter norm govern loss dynamics primarily through their ratio, the effective learning rate (ELR). When ELR is matched across runs, their loss trajectories collapse throughout training despite substantially different LRs and parameter norms. Across optimizers, architectures, datasets, and model scales, mean collapse errors are typically a few x 10^-3, below the seed-to-seed variation measured in a representative configuration. Systematic ablations identify normalization design and the timescale of LR-norm variation as key determinants of collapse precision. Controlled interventions further show that weight decay and Hyperball shape loss dynamics primarily through the ELR schedules they induce. Replacing LR with ELR enables a fitted functional scaling law (FSL) to transfer across norm-control methods. The resulting ELR-based FSL also explains delayed acceleration, a recurring effect of norm control. Together, these results establish ELR as a common coordinate linking LR scheduling, norm control, and loss dynamics.
发表机构
- Peking University(北京大学)
- Ant Group(蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。