发表机构
California Institute of Technology; Tsinghua University; University of Wisconsin - Madison(加州理工学院; 清华大学; 威斯康星大学麦迪逊分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AutoLoCo提出自适应同步框架,通过动态调整局部更新间隔并修正外部优化器,在保持训练性能的同时降低27%的通信频率。
AI 中文摘要
大型语言模型(LLM)的预训练日益在多个数据中心之间进行。随着训练扩展到更大数量的加速器,用于计算的时间比例下降,而用于通信的时间比例上升。因此,频繁的同步成为一个日益严重的瓶颈。局部更新方法通过允许工作节点在同步之间执行多次优化器步骤来减少这一成本。大多数局部更新方法在训练前设定同步之间的局部优化器步骤数,并在整个运行过程中保持该间隔固定。然而,最佳间隔可能在训练过程中发生变化。如果间隔和优化器适应当前训练状态,则可以在保持训练性能的同时降低通信频率。在这项工作中,我们引入了AutoLoCo,一个用于减少LLM训练中通信的自适应训练框架。它利用标量训练统计量自适应地调整局部间隔,并修正每次外部更新。我们的方法基于两个观察:1)合适的局部间隔在训练阶段之间有所不同,2)改变每个间隔的内部步骤数会与未改变的外部优化器产生不匹配,需要对外部更新进行修正。我们通过使用累积的内部学习率对外部优化器的动量和学习率进行修正来优化这种不匹配。我们在通信约束下的实验表明,与DiLoCo相比,AutoLoCo将通信频率降低了27%,同时保持了训练性能。
英文摘要
The pre-training of Large Language Models (LLMs) is increasingly conducted across multiple data centers. As training scales to a larger number of accelerators, the fraction of time spent on computation decreases, while the fraction spent on communication increases. Therefore, frequent synchronization becomes a growing bottleneck. Local update methods reduce this cost by allowing workers to perform several optimizer steps between synchronizations. Most local update methods set the number of local optimizer steps between synchronizations before training and keep this interval fixed throughout the run. However, the best interval can change during the entire train process. If the interval and optimizer are adapted to the current training state, the communication frequency is reduced while maintaining the training performance. In this work, we introduce AutoLoCo, an adaptive training framework to reduce communication in LLM training. It adapts the local interval using scalar training statistics and corrects each outer update. Our method is motivated by two observations: 1) the appropriate local interval varies across training stages, and 2) changing the number of inner steps per interval creates a mismatch with an unchanged outer optimizer, requiring a correction to the outer update. We optimize this mismatch by correction of the outer optimizer for the momentum and the learning rate using the accumulated inner learning rate. Our experiments under communication constraints demonstrate that AutoLoCo reduces communication frequency by 27% relative to DiLoCo while maintaining training performance.