发表机构
University of Illinois at Urbana-Champaign; IBM Research(伊利诺伊大学厄巴纳-香槟分校; IBM研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
TRANSIT是一个透明的规模缩减框架,通过用户空间介入层和零拷贝路径扩展GPU内存,在减少GPU数量的同时保持高效的多节点LLM训练,显著提升吞吐量并降低网络流量。
AI 中文摘要
TRANSIT是一个透明的规模缩减框架,旨在使用更少的GPU进行多节点模型训练,同时通过在分布式训练期间透明地将CPU DRAM用作GPU内存的扩展来保持训练效率。它通过用户空间介入层实现这一点,无需修改应用程序、训练框架、集群调度器、设备驱动程序或操作系统。此外,TRANSIT通过利用CPU-GPU传输的零拷贝数据路径实现了更高的效率。我们在密集模型和MoE模型上,在多达64个NVIDIA H100 GPU和多种并行配置下,通过RoCE网络评估了TRANSIT。我们的评估表明,TRANSIT能够:(a) 超越最先进的框架管理卸载技术,每GPU吞吐量分别比TorchTitan、ZeRO-Offload和ZeRO-Infinity高出高达68%、59%和42%;(b) 在保持基线每GPU吞吐量超过90%的同时,使用减少50%的GPU进行训练;(c) 将每节点网络流量降低高达33%;以及(d) 在通信受限的设置中,将每GPU吞吐量提高高达35%。
英文摘要
TRANSIT is a transparent scale-in framework to enable multi-node model training on fewer GPUs while maintaining training efficiency by transparently leveraging CPU DRAM as an extension of GPU memory during distributed training. It achieves this through a user-space interposition layer, requiring no modifications to the application, training framework, cluster scheduler, device driver, or operating system. Furthermore, TRANSIT achieves higher efficiency by leveraging a zero-copy data path for CPU-GPU transfers. We evaluate TRANSIT on dense and MoE models across scales up to 64 NVIDIA H100 GPUs and multiple parallelism configurations over a RoCE network. Our evaluation shows that TRANSIT can: (a) outperform state-of-the-art framework-managed offloading techniques, achieving up to 68%, 59%, and 42% higher per-GPU throughput than TorchTitan, ZeRO-Offload, and ZeRO-Infinity, respectively, (b) enables training with 50% fewer GPUs while maintaining over 90% of baseline per-GPU throughput, (c) lower per-node network traffic by up to 33%, and (d) improve per-GPU throughput by up to 35% in communication-bound settings.
Comments15 pages, 12 figures