arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06506cs.DCcs.PF

共享集体通信网络:深度学习训练中的两种存储开销

Sharing a Fabric with Collective Communication: Two Storage Penalties in Deep Learning Training

Chen Wang, Wenzhao Wu, Hyojin Kim, Jae-Sung Yeom

首次发表
浏览论文内容

中文总结 AI 辅助

本文发现HPC系统中深度学习训练共享网络导致存储I/O与集体通信相互干扰,产生DataLoader停滞和网络争用两种开销,提出节点本地NVMe暂存方案DYAD,实现最高7.4倍加速。

中文摘要 AI 辅助

在HPC系统上进行分布式深度学习训练时,NCCL/RCCL集体通信与并行文件系统I/O通常共享同一网络结构。通过在Slingshot-11系统上使用真实的GNN训练工作负载,我们展示了这种共享会带来两种不同的开销。主要开销是重尾的DataLoader停滞:稳态下典型的DataLoader等待仅为15毫秒,但在28%的Lustre迭代和12%的VAST迭代中,等待时间会飙升至数秒。次要开销是集体通信上的流量类别争用:在隔离基准测试中,Lustre I/O使all-reduce操作停滞高达145倍。这两种开销源于不同的机制。I/O停滞延迟影响任何穿越共享网络的存储路径,而all-reduce网络争用仅在存储和集体通信共享同一流量类别时发生。它们的共同根本原因是存储I/O穿越了共享网络。这项工作表明,通过DYAD(我们的代码可在以下https URL公开获取)进行节点本地NVMe暂存,通过将存储I/O保持在共享路径之外,消除了这两种影响。在整个训练周期中,DYAD相比直接Lustre读取实现了7.4倍加速,相比VAST实现了1.06倍加速。到第二个周期,一旦本地缓存完全预热,DataLoader停滞被完全消除,使DYAD相比VAST达到1.31倍加速。

英文摘要

Distributed DL training on HPC systems often shares one network fabric between NCCL/RCCL collective communication and parallel-filesystem I/O. Using a real GNN training workload on a Slingshot-11 system, we show that this sharing imposes two distinct costs. The primary cost is heavy-tailed DataLoader stalls: the typical DataLoader wait is just 15 ms at steady state, yet spikes to multiple seconds in 28% of Lustre iterations and 12% of VAST iterations. The secondary cost is traffic-class contention on collective communication: Lustre I/O stalls the all-reduce by up to 145$\times$ in an isolated benchmark. The two costs arise from different mechanisms. I/O stall latency affects any storage path that traverses the shared fabric, whereas all-reduce network contention occurs only when storage and collective communication share the same traffic class. Their common root cause is that storage I/O traverses the shared fabric. This work shows that node-local NVMe staging via DYAD (Our code is publicly available at https://github.com/flux-framework/dyad) eliminates both effects by keeping storage I/O off that path. Across a full training epoch, DYAD achieves a 7.4 times speedup over direct Lustre reads and a 1.06 times speedup over VAST. By the second epoch, once the local cache is fully warmed, DataLoader stalls are eliminated entirely, allowing DYAD to reach a 1.31 times speedup over VAST.

发表机构

  • CCDS NTU Singapore(南洋理工大学计算通信与数据科学中心)
  • CASC LLNL(劳伦斯利弗莫尔国家实验室计算、算法和系统中心)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑