发表机构
The Hong Kong University of Science and Technology (Guangzhou); IDEA Research; DataArcTech Ltd.(香港科技大学(广州); IDEA研究院; DataArcTech有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
LazyTrain是面向大语言模型训练的优化层,将相关问题建模为混合整数调度问题,结合Hybrid 8位算子,在H800、RTX 3090实验中提升算力、批次大小与准确率。
AI 中文摘要
在有限硬件条件下训练大语言模型日益成为GPU算力、主机内存、PCIe传输与存储带宽之间的调度问题。现有卸载系统可减少GPU驻留内存,MegaTrain表明CPU主导的层流式执行器可在单GPU上训练大模型,但固定检查点设置与放置启发式方法仍会在关键路径上暴露通信开销。我们提出LazyTrain,这是层流式执行器之上的优化层。LazyTrain将检查点选择、激活值放置、重计算以及CPU-GPU-NVMe通信重叠问题建模为混合整数调度问题,在训练过程中执行求解得到的策略。它还将8位优化器状态与快速梯度裁剪耦合为单一的Hybrid 8位算子:状态压缩可减少优化器状态内存,而快速裁剪可抵消额外的CPU端更新开销。在H800上针对Qwen2.5-3B至Qwen3.6-27B的实验中,LazyTrain相比匹配基准运行将持续TFLOPS提升约1.24倍;RTX 3090实验在各模型规模下均将最大可行批次大小提升1。在主要的Qwen3.6-27B H800 MetaMathQA运行中,LazyTrain在批次大小72时达到219.95 TFLOPS、1361 tokens/s,峰值GPU内存为68.84GB,在完整评估拆分上获得95.42%的精确匹配准确率。源代码可在该https URL获取。
英文摘要
Training large language models on limited hardware is increasingly a scheduling problem across GPU compute, host memory, PCIe transfer, and storage bandwidth. Existing offloading systems reduce GPU residency, and MegaTrain shows that a CPU-master layer-streaming executor can train large models on a single GPU, but fixed checkpointing and placement heuristics still leave communication exposed on the critical path. We propose LazyTrain, an optimization layer over a layer-streaming executor. LazyTrain formulates checkpoint selection, activation placement, recomputation, and CPU-GPU-NVMe communication overlap as a mixed-integer scheduling problem, then executes the solved policy during training. It further couples 8-bit optimizer states with fast gradient clipping as a single Hybrid 8-bit operator: state compression reduces optimizer-state memory, while fast clipping counteracts the additional CPU-side update overhead. Across H800 experiments from Qwen2.5-3B to Qwen3.6-27B, LazyTrain improves sustained TFLOPS over matched baselines runs by approximately 1.24$\times$; RTX 3090 experiments likewise increase the maximum feasible batch size by one at each model scale. In the primary Qwen3.6-27B H800 MetaMathQA run, LazyTrain reaches 219.95 TFLOPS and 1361 tokens/s at batch size 72, peaks at 68.84\,GB of GPU memory, and obtains 95.42\% exact-match accuracy on the full evaluation split. The source code is available at https://github.com/DataArcTech/LazyTrain.
Comments18 pages, 8 figures