arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

拓扑在机器学习训练性能中的作用

On Topology's Role in ML Training Performance

Sarah McClure, Tegan Wilson, Brad Karp, Michael Mitzenmacher, Sylvia Ratnasamy, Scott Shenker, Minlan Yu

arXiv 2608.01707首次发表:更新:

AI 中文总结

该研究分析了胖树Clos、圆环等拓扑对机器学习训练中集合通信操作性能的影响,发现Clos在多数场景下集合完成时间更优,且弹性、灵活性更好,不存在全场景占优的拓扑。

AI 中文摘要

现代机器学习训练工作负载运行在由计算加速器组成的大规模网络上,这类系统中常用的网络拓扑通常是胖树Clos和圆环这两种基本拓扑的变体。本文推导了分析结果,阐明了拓扑选择如何影响支撑现代机器学习工作负载的少量集合通信操作的可实现性能。我们还考虑了纳入网络故障、作业放置策略等额外因素时这些结果的变化。总体而言,不存在在所有场景都占优的拓扑,但Clos在多数情况下能实现更优的集合完成时间,且在弹性和灵活性方面具有优势。

英文摘要

Modern machine learning training workloads run on large-scale networks of compute accelerators. The networks commonly deployed in these systems are typically variations of two basic topologies: the fat-tree Clos and the torus. In this paper, we derive analytical results the elucidate how the choice of topology shapes achievable performance for the small set of collective communication operations that underlies modern machine learning workloads. We also consider how these results change when we include additional factors such as network failures and job placement strategies. Overall, we find that one topology does not dominate in all cases, but that the Clos achieves better collective completion time in most cases and provides benefits in resilience and flexibility.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑