分布式大语言模型系统的集体通信:规划、运行时适配与计算协调
Collective Communication for Distributed LLM Systems: Planning, Runtime Adaptation, and Computation Coordination
浏览论文内容
中文总结 AI 辅助
本文提出以集体为中心的分类法,将分布式LLM系统集体通信进展分为规划、执行适配、计算通信协调三层,探讨相关开放挑战与未来机遇。
中文摘要 AI 辅助
分布式大语言模型(LLM)系统日益依赖集体通信原语,如AllReduce(AR)、ReduceScatter(RS)、AllGather(AG)和AlltoAll(A2A)。在现代LLM训练与服务集群中,异构GPU互连、多NIC网络、混合并行策略、低延迟推理请求及高吞吐量训练流水线,推动集体通信的规划、执行与重叠方式愈发多样化。本文提出一种以集体为中心的教程式分类法,将最新进展分为三层:通信规划层生成感知拓扑的集体调度方案;通信执行与适配层将这些调度方案映射到真实集群的GPU运行时与硬件;计算-通信协调层将集体优化转化为端到端训练与推理的收益。本文还探讨了分布式LLM系统中集体通信的开放挑战与未来机遇。
英文摘要
Distributed large language model (LLM) systems increasingly rely on collective communication primitives such as AllReduce (AR), ReduceScatter (RS), AllGather (AG), and AlltoAll (A2A). In modern LLM training and serving clusters, heterogeneous GPU interconnects, multi-NIC networking, mixed parallelism strategies, low-latency inference requests, and high-throughput training pipelines have motivated increasingly diverse ways to plan, execute, and overlap collective communication. This paper presents a tutorial-style, collective-centric taxonomy for collective communication. We organize recent advances into three layers: communication planning, which generates topology-aware collective schedules; communication execution and adaptation, which maps these schedules onto GPU runtimes and hardware in real clusters; and computation-communication coordination, which turns collective optimization into end-to-end training and inference benefits. We further discuss open challenges and future opportunities for collective communication in distributed LLM systems.