发表机构
The University of Texas at Austin(德克萨斯大学奥斯汀分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Leto利用幸存硬件实现原地恢复,通过保留模型状态和影子预初始化,显著加速LLM训练故障恢复,提升有效训练时间。
AI 中文摘要
硬件可操作故障(HOF)会中断大语言模型(LLM)的训练,但允许在同一硬件上进行恢复,无需重置、维修或更换。然而,现有的恢复系统仍然会重新加载检查点、重新计算丢失的进度并重建进程状态,导致本可以继续训练的GPU闲置。我们提出了Leto,一个容错训练系统,利用幸存硬件实现高效的原地恢复。我们的关键见解是,恢复训练所需的状态可以在活动训练进程之外保留或准备,同时保持在相同的硬件上。Leto保留工作模型状态和可复用的进程状态,并在影子训练器中预初始化剩余状态。我们设计了双层擦除保护和分块级事务更新,以保持保留的模型状态可恢复且一致,并在活动训练需要其GPU内存时回收影子状态。在6-GPU和72-GPU的NVIDIA A100集群上的评估显示,Leto的恢复速度比性能最佳的检查点基线快3.6至6.5倍,并将有效训练时间提高了多达13.7个百分点。大规模模拟显示,在131,072个GPU的集群上,有效训练时间超过95%。
英文摘要
Hardware-operable failures (HOFs) interrupt large language model (LLM) training but permit recovery on the same hardware without reset, repair, or replacement. Existing recovery systems nevertheless reload checkpoints, recompute lost progress, and rebuild process state, idling GPUs that could otherwise continue training. We present Leto, a fault-tolerant training system that leverages surviving hardware to enable efficient in-place recovery. Our key insight is that the state needed to resume training can be retained or prepared outside the active training process while remaining on the same hardware. Leto retains the working model state and the reusable process state, and preinitializes the remaining state in a shadow trainer. We devise two-tier erasure protection and chunk-level transactional updates to keep the retained model state recoverable and consistent, and reclaim the shadow state when active training needs its GPU memory. Evaluation on 6- and 72-GPU NVIDIA A100 clusters shows that Leto recovers 3.6--6.5$\times$ faster than the best-performing checkpointing baselines and improves productive training time by up to 13.7 percentage points. Large-scale simulation shows over 95% productive training time on a 131,072-GPU cluster.