发表机构
University of Turin(都灵大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本工作受联邦学习通信原理启发,提出FL+FSDP和FL+HSDP两种混合算法,将分片数据并行解耦为松散耦合的联邦组,在512个A100 GPU上预训练Llama3.1 8B,实现高达8.04倍数据处理加速和4.48倍困惑度降低。
AI 中文摘要
人工智能模型与高性能计算系统的共生扩展不断在其收敛过程中产生算法挑战。基础模型(FMs)是一个关键例子,需要在数千个尖端GPU上进行长达数月的训练。分片数据并行(DP)是通过将数据和模型分割到多个GPU上来加速此类计算的主导策略。然而,当大规模部署时,尤其是在具有异构性能的多层互连上,它会产生高昂的通信开销。受联邦学习(FL)高效通信原理的启发,本工作引入了两种混合算法——FL+FSDP和FL+HSDP——将分片DP与FedAvg风格的聚合交错进行。这种方法将大型DP部署解耦为更小的、松散耦合的联邦组,在保持全局批大小受组大小限制的同时,最小化组间流量。对通信成本的正式分析和实验验证证明了它们的可扩展性和灵活性。在512个A100 GPU上进行的Llama3.1 8B预训练表明,在相同超参数下,FL+FSDP和FL+HSDP相比对应方法实现了高达8.04倍的数据处理加速和4.48倍的评估困惑度降低,展示了卓越的计算效率和改进的模型质量。这些特性源于通信开销的减少以及全局批大小相对于联邦组大小的有界增长。
英文摘要
The symbiotic scaling of artificial intelligence models and high-performance computing systems continually creates algorithmic challenges in their convergence. Foundation models (FMs) are a crucial example, requiring months-long training on thousands of cutting-edge GPUs. Sharded data parallelism (DP) is the dominant strategy to accelerate such computations by splitting data and models across multiple GPUs. However, it incurs prohibitive communication overhead when deployed at scale, particularly on multi-tier interconnects with heterogeneous performance. Inspired by the efficient communication principles of federated learning (FL), this work introduces two hybrid algorithms - FL+FSDP and FL+HSDP - interleaving sharded DP with FedAvg-style aggregations. Such approaches decouple large DP deployments into smaller, loosely-coupled federation groups, requiring minimal inter-group traffic while keeping the global batch size bounded by the groups' size. Formal analysis of communication costs and experimental validation prove their scalability and flexibility. A Llama3.1 8B pre-training on 512 A100 GPUs shows that, under identical hyperparameters, FL+FSDP and FL+HSDP achieve up to 8.04 faster data processing and 4.48 lower evaluation perplexity than their counterparts, demonstrating superior computational efficiency and improved model quality. These properties stem from reduced communication overhead and the bounded growth of the global batch size relative to the federation group size.
Journal refEuro-Par 2026: Parallel Processing - 32nd European Conference on Parallel and Distributed Processing, Pisa, Italy, August 24-28, 2026, Proceedings, Part II
DOI:10.1007/978-3-032-35251-4_30