发表机构
Meta Platforms, Inc.(Meta平台公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大规模推荐训练中生命周期开销浪费加速器容量的问题,提出基于有效训练时间(ETT%)框架的系列优化,涵盖通信消除、缓存复用等,使ETT%平均提升15.5%,集群范围从80%提升至90%以上。
AI 中文摘要
生命周期开销悄然消耗着大规模推荐训练集群中的加速器容量。我们最大的推荐工作负载每天在数千块GPU上处理数百亿个训练样本。在本工作之前,其端到端挂钟时间中仅有50-60%用于在新数据上推进训练。我们针对这一生命周期开销进行了集群规模的研究,并提出了一系列涵盖整个训练栈的优化措施。我们使用有效训练时间(ETT%)作为操作框架,对丢失的时间进行测量,将其定位到独立拥有的基础设施组件,并暴露跨作业重启的重复工作。该分析指导了诸如在训练器初始化期间消除通信和进行流水线重叠等优化;动态形状处理、自动调优剪枝和可复用的PyTorch 2编译缓存;异步检查点;独立的模型发布;以及降低恢复成本。我们在代表性模型上评估了这些优化,并测量了它们在我们训练集群中的影响。ETT%在每项基准测试中均有所提升,平均提升15.5%,在我们最大的工作负载上达到85%。部署后,集群范围的ETT%从约80%上升至90%以上。
英文摘要
Lifecycle overhead silently consumes accelerator capacity across large-scale recommendation training fleets. Our largest recommendation workloads process tens of billions train- ing examples per day on thousands of GPUs. Before this work, only 50-60% of their end-to-end wall time advanced training on new data. We present a fleet-scale study of this lifecycle overhead and a set of optimizations spanning the full training stack. We use Effective Training Time (ETT%) as an operational framework to instrument lost time, localize it to independently owned infrastructure components, and expose work repeated across job restarts. This analysis guides optimizations like communication elimination and pipeline overlap during trainer initialization; dynamic-shape handling, autotuning pruning, and reusable Py- Torch 2 compilation caches; asynchronous checkpointing; stan- dalone model publishing; and reductions in recovery cost. We evaluate the optimizations on representative models and measure their impacts in our training fleet. ETT% improves on every benchmark, by 15.5% on average, and reaches 85% on our largest workload. Fleet-wide ETT% rose from about 80% to above 90% after deployment.