arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

调整集体模式以缓解共享AI集群中的拥塞

Tuning Collective Patterns to Alleviate Congestion in Shared AI Clusters

Eashan Gupta, Yongzhou Chen, Apoorve Mohan, Pavlos Maniotis, Abdullah Kayi, Radhika Mittal

arXiv 2609.04417首次发表:更新:

发表机构

UIUC; Meta; IBM Research(伊利诺伊大学厄巴纳-香槟分校; Meta; IBM研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对共享AI集群的外部拥塞问题,提出基于NCCL垫片层的REACT系统,通过调整通信集体模式缓解拥塞,在实测和模拟中均实现显著性能提升。

AI 中文摘要

分布式AI训练涉及多对GPU节点之间的多轮数据交换,即使某一条流因拥塞而变慢,也可能导致整个通信轮次变慢。当前AI集群中避免拥塞的方法要么假设对整个工作负载有全局控制权(例如协调所有作业的调度),要么假设存在基础设施支持(例如交换机中的自适应路由),因此不适用于共享云环境——在该环境中,属于一个用户的AI作业可能会面临来自其他用户作业或自身无法控制的背景流量的外部拥塞。本文构建了名为REACT的系统,它会根据拥塞情况调整GPU节点之间数据交换的重复模式(即通信集体)。REACT工作在应用层(通信库层),它利用现成的流统计信息在运行时检测拥塞,并调整集体模式以缓解拥塞——在保留信息交换语义的同时改变入射流的集合(例如在AllReduce树中选择哪个节点聚合数据)。REACT不需要底层网络基础设施的显式支持,单个用户可在共享云环境中单方面部署。我们将REACT原型化为NCCL的一个垫片层,并在一个共享的学术GPU集群上对其进行评估:在网络拥塞下,启用REACT可将通信性能(算法带宽)提升13%-38%。我们在一系列拥塞场景下的模拟进一步显示,性能提升最高可达75%,凸显了我们方法的有效性。

英文摘要

Distributed AI training involves recurring rounds of data exchange between multiple pairs of GPU nodes. Slowdown in even one flow due to congestion can cause the entire communication round to slowdown. Current approaches for evading congestion in AI clusters assume global control over the entire workload (e.g. coordinating the schedule of all jobs) or assume infrastructural support (e.g. adaptive routing in switches). They are thus ill-suited in a shared cloud setting where AI jobs belonging to one user can face external congestion from other users' jobs or background traffic beyond its own control. In this paper, we build a system, REACT, that tunes the recurring pattern of data exchange between GPU nodes (known as communication collectives) in response to congestion. REACT works at the application (communication library) layer, where it detects congestion at runtime using readily available flow stats, and tunes the collective pattern to alleviate congestion - changing the set of incident flows while retaining the semantics of information exchange (e.g. selecting which node aggregates data in an AllReduce tree). REACT requires no explicit support from the underlying network infrastructure and can be unilaterally deployed by individual users in a shared cloud setting. We prototype REACT as a shim layer over NCCL, and evaluate it on a shared academic GPU cluster - enabling REACT improves communication performance (algorithm bandwidth) by 13%-38% under network congestion. Our simulations across a range of congestion scenarios further reveal up to 75% performance improvement, highlighting the effectiveness of our approach.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑