发表机构
Meta(Meta)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究从凸优化视角推导了适用于通用优化器和模型架构的损失联合表征,得到闭式最优批量大小调度与联合缩放律,其性能优于固定批量大小基线,凸显动态批量大小调度在大语言模型训练中的重要性。
AI 中文摘要
现代深度学习通常在整个训练过程中保持批量大小固定,从而忽略了学习率与批量大小对训练动态的联合影响。本文从凸优化视角研究深度学习动态,推导了适用于通用优化器和模型架构的、关于学习率和批量大小调度的损失联合表征。该表征可针对任意给定学习率调度得到闭式最优批量大小调度,进而得到联合缩放律,其性能始终优于固定批量大小基线,凸显了动态批量大小调度在大语言模型训练中的重要性。
英文摘要
Modern deep learning typically keeps the batch size static throughout training, thus overlooking the joint effect of learning rate and batch size on the training dynamics. In this paper, we study the deep learning dynamics through the lens of convex optimization and derive a joint characterization of loss in terms of both schedules, applicable to general optimizers and model architectures. This characterization yields a closed-form optimal batch size schedule for any prescribed learning rate schedule, and further leads to joint scaling laws that consistently outperform static batch size baselines, highlighting the significance of dynamic batch size schedule in large language model training.