arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34162cs.DC

SlideDP:跨多GPU扩展宿主机驻留的大语言模型微调

SlideDP: Scaling Host-Resident LLM Fine-Tuning Across Multiple GPUs

Ruijia Yang, Shiyuan Lin, Yulong Ao, Zhiyu Li, Yingli Zhao, Xianduo Li, Yonghua Lin, Zeyi Wen

首次发表
浏览论文内容

中文总结 AI 辅助

SlideDP提出同步数据并行运行时,通过宿主机状态权威、通信解耦和流水线化,在共享宿主机多GPU上实现高效大语言模型微调,吞吐量优于现有系统。

中文摘要 AI 辅助

宿主机驻留的层流式传输使得全参数大语言模型微调能够超越GPU内存的限制,但数据并行等级会竞争共享的宿主机资源。复制传输会放大通信量,而强扩展可能在计算窗口缩小时暴露宿主机工作负载。我们提出了SlideDP,一个面向共享宿主机多GPU系统的同步数据并行运行时。它维护一个权威的宿主机状态,将通信路径与状态布局解耦,并在等级和块之间流水线化参数传递、梯度聚合和CPU更新。一个解析的步时间模型刻画了资源瓶颈和流水线暴露;运行时测量在GPU内存预算下指导通信、分块和激活策略。在匹配批次的扫描中,SlideDP相比SlideFormer、MegaTrain和ZeRO-Offload实现了几何平均吞吐比1.46-2.64倍。在四块H100上,SlideDP在较小批次下接近GPU驻留的FSDP2对Qwen3-14B的吞吐量。使用更大批次时,它每步处理超过100万令牌,并超出FSDP2实测峰值吞吐量11.2%。此外,它支持同一模型的256K令牌序列,并在四块RTX 4090 GPU上微调Qwen2.5-72B。项目页面:此https URL。

英文摘要

Host-resident layer streaming enables full-parameter LLM fine-tuning beyond GPU memory, but data-parallel ranks compete for shared host resources. Replicated transfers amplify traffic, while strong scaling can expose host work as computation windows shrink. We present SlideDP, a synchronous data-parallel runtime for shared-host multi-GPU systems. It maintains one authoritative host state, decouples communication routes from state layout, and pipelines parameter delivery, gradient aggregation, and CPU updates across ranks and chunks. An analytical step-time model characterizes resource bottlenecks and pipeline exposure; runtime measurements guide communication, chunking, and activation policies under a GPU memory budget. In matched-batch sweeps, SlideDP achieves geometric-mean throughput ratios of 1.46-2.64$\times$ over SlideFormer, MegaTrain, and ZeRO-Offload. On four H100s, SlideDP approaches GPU-resident FSDP2 throughput for Qwen3-14B at a smaller batch size. With a larger batch, it processes over 1M tokens per step and exceeds FSDP2's measured peak throughput by 11.2%. Separately, it supports 256K-token sequences for the same model and fine-tunes Qwen2.5-72B on four RTX 4090 GPUs. Project page: https://github.com/RegiaYoung/SlideDP.

发表机构

  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
  • Beijing Academy of Artificial Intelligence(北京智源人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑