arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.09657cs.DC

快速且内存高效的混合XPU计算卸载训练框架

Fast and Memory Efficient Offload Training Framework with Hybrid XPU Computation

Zhiyi Yao, Zuning Liang, Yuedong Xu, Jin Zhao, Jessie Hui Wang, Tong Li

首次发表
浏览论文内容

中文总结 AI 辅助

提出MemFerry框架,利用GPU直接主机访问实现混合计算,通过调度和影子模型优化卸载,显著提升训练速度和模型规模。

中文摘要 AI 辅助

随着深度学习模型规模的不断增大,GPU内存往往在训练过程中变得不足。一种突出的方法是ZeRO-Offload,它将优化器状态移至CPU内存,并使用CPU执行参数更新。然而,ZeRO-Offload的缺陷包括GPU利用率低、通信与计算重叠不完善以及卸载不灵活。在本文中,我们利用GPU上的直接主机访问(DHA)技术,能够在CPU内存中计算数据,形成一种新颖的GPU与DHA混合计算模式。我们设计并实现了MemFerry,它由一个执行调度器和一个影子模型组成。调度器策略性地选择参数层进行DHA计算,并同时将剩余参数传输到GPU内存,以缩短前向传播时间,并进一步将DHA参数加载到GPU内存以减少反向传播时间。影子模型为分别存储在GPU和CPU内存中的参数分区提供了统一的内存抽象。为了进一步减少GPU内存使用,我们提出了MemFerry及其动态规划算法,通过DHA将梯度卸载到CPU内存。我们进一步将MemFerry扩展到新兴的扩展域,提出了ScaleUp-MemFerry,它利用原本未充分利用的加速器互连带宽,通过自适应多路径传输协助主机到加速器的数据移动。我们的实验表明,与单GPU上的ZeRO-Offload相比,MemFerry的训练速度最高可提升1.68倍,并且能够训练大1.52倍的模型;在扩展到8个GPU的数据并行时,训练速度至少提升28.1%。我们进一步将设计扩展到华为CloudMatrix384扩展节点,该节点最多包含8个NPU,我们的ScaleUp-MemFerry相比DeepSpeed,端到端迭代时间最多减少20.7%。

英文摘要

With the ever-growing size of deep learning models, GPU memory is prone to being insufficient during training. A prominent approach is ZeRO-Offload, which moves the optimizer states to CPU memory and performs parameter update using CPU. However, the deficiencies of ZeRO-Offload include low GPU utilization, imperfect overlapping of communication and computation, and inflexible offloading. In this paper, we leverage Direct Host Access (DHA) on the GPU that can compute data in CPU memory, forming a novel hybrid on-GPU and DHA. We design and implement MemFerry consisting of an execution scheduler and a shadow model. The scheduler strategically chooses layers of parameters for DHA computation and transmits the remaining parameters to GPU memory simultaneously to shorten forward propagation time, and further loads DHA parameters to GPU memory to reduce backward propagation time. The shadow model presents a unified memory abstraction for the parameter partitions stored separately in GPU and CPU memories. To further reduce GPU memory usage, we present MemFerry along with its dynamic programming algorithm that offloads gradients to CPU memory via DHA. We further extend MemFerry to emerging scale-up domains with ScaleUp-MemFerry, which exploits otherwise underutilized accelerator interconnect bandwidth to assist host-to-accelerator data movement through adaptive multi-path transfer. Our experiments show that \system trains up to $1.68\times$ faster and MemFerry can train $1.52\times$ larger model compared to ZeRO-Offload on a single GPU, and increase training speed by at least $28.1\%$ when scaling to data parallelism on 8 GPUs. We further extend the design to a Huawei CloudMatrix384 scale-Up node with up to 8 NPUs, and our ScaleUp-MemFerry reduces the end-to-end iteration time by up to $20.7\%$ over DeepSpeed.

发表机构

  • College of Future Information Technology, Fudan University(复旦大学未来信息学院)
  • Institute for Network Sciences and Cyberspace, Tsinghua University(清华大学网络科学与网络空间研究院)
  • School of Information, Renmin University of China(中国人民大学信息学院)
  • College of Computer Science and Artificial Intelligence, Fudan University(复旦大学计算机科学与人工智能学院)

机构由 AI 辅助整理,请以论文原文为准。

↑