多租户环境下大规模AI训练的可扩展性与性能表征
Characterizing the Scalability and Performance of Large-Scale AI Training Under Multi-Tenancy
- University of Trento(特伦托大学)
- Sapienza University of Rome(罗马第一大学)
- NVIDIA(英伟达)
- ENEA(意大利国家新能源与可再生能源机构)
- Forschungszentrum Jülich GmbH(于利希研究中心有限公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究针对多租户场景下的大规模AI训练,设计基准套件评估5种并行策略在6类超级计算机集群的性能,量化通信开销,揭示并发作业干扰,系统表征分布式AI训练的可扩展性与执行效率。
AI中文摘要:
表征现代高性能计算(HPC)系统上的AI工作负载性能,需同时理解其孤立运行时的可扩展性与并发执行下的行为,但并行策略、网络拥塞、计算能力与互连技术之间的相互作用仍知之甚少。本研究探究最多2400个GPU的AI模型的性能与可扩展性,通过在多种分配方案下评估规模提升、横向扩展及机架级配置,量化通信开销及其在不同互连方式下的影响;还通过设计真实噪声模型,研究多个并发训练作业间的相互干扰。本研究设计了一套AI模型基准测试套件,用于评估Alps、Leonardo、LUMI、JUPITER、NVL72 GB300及DGX A100等不同超级计算机集群上五种不同并行策略的性能,系统表征了分布式AI训练的可扩展性与执行效率,为真实多租户场景下的性能行为提供关键见解。
英文摘要:
Characterising AI workload performance on modern HPC systems requires understanding both their scalability in isolation and their behaviour under concurrent execution. However, the interplay among parallelisation strategies, network congestion, compute capability, and interconnect technologies remains poorly understood. This work investigates the performance and scalability of AI models up to 2400 GPUs. We quantify the communication overheads and their impact across different interconnects by evaluating scale-up, scale-out, and rack-scale configurations under multiple allocation schemes. Finally, we study how multiple concurrent training jobs interfere with each other by designing a realistic noise model. We design a benchmark suite of AI models to evaluate the performance of five distinct parallelisation strategies across different supercomputing clusters, including Alps, Leonardo, LUMI, JUPITER, NVL72 GB300, and DGX A100. Our work provides a systematic characterization of the scalability and execution efficiency of distributed AI training, while offering key insights into performance behavior under realistic multi-tenant scenarios.