AI 中文总结
针对分布式训练中环形全归约通信开销大的问题,提出CERAR协议,通过线性编码实现近似梯度聚合,降低通信和存储开销,并在多GPU集群上验证了低带宽场景下的有效性。
AI 中文摘要
环形全归约(Ring All-Reduce)被广泛用于大规模分布式训练中,以精确计算梯度总和。对于具有 $N$ 个工作节点的系统,其归一化的每工作节点通信量为 $2(N-1)/N$。在本工作中,我们提出了一种通信高效的环形全归约(CERAR)协议,用于“近似”梯度聚合。CERAR 将每个本地梯度划分为 $c$ 个分量,并关键依赖于线性编码和解码操作。它在 $N$ 工作节点环上执行 $L=c+N-2$ 轮通信,得到归一化通信速率 $1+(N-2)/c$ 和归一化存储 $1+N/c$。我们给出了一种显式的基于范德蒙德(Vandermonde)的构造,其近似误差可以随着 $c$ 的增加而任意接近零,同时通信和存储速率趋近于一。然而,这一极限是通过病态编码矩阵实现的。因此,我们给出了一种乘法扰动构造,其误差为 $O(\epsilon)$,而相关的条件数为 $O(\epsilon^{-r_\star})$,其中 $r_\star=\lceil c/N\rceil-1$。这自然引出了一个条件数约束的优化公式,用于权衡相互竞争的目标并获得数值稳定的实际设计。我们在具有低带宽和高带宽互连的多 GPU 集群上进行了实验。我们的结果表明,在低带宽设置下,即使参数长度适中,也有明显的优势。我们预期在高带宽设置下,对于参数长度大得多的实验,也会有相应的改进。
英文摘要
Ring All-Reduce is a widely used protocol for aggregating gradients across workers in large-scale distributed training. For a system with $N$ workers, its normalized per-worker communication is $2(N-1)/N$. In this work we present a communication-efficient Ring All-Reduce (CERAR) protocol for ``approximate'' gradient aggregation. CERAR partitions each local gradient into $c$ components and relies on linear encoding and decoding operations. It performs $L=c+N-2$ communication rounds over an $N$-worker ring, yielding normalized communication rate $1+(N-2)/c$ and normalized storage rate $1+N/c$. We present an explicit Vandermonde matrix based construction whose approximation error can be made arbitrarily small, with communication and storage rates approaching one as $c$ increases. However, this comes at the cost of ill-conditioned encoding matrices. We therefore give an alternate construction whose error is $O(ε)$, while the relevant condition numbers are $O(ε^{-r_\star})$, where $r_\star=\lceil c/N\rceil-1$. This motivates a condition-number-constrained optimization formulation for obtaining numerically stable practical designs. Our experiments show that, compared with the production-standard NCCL on NVIDIA A100 PCIe GPU clusters with low-bandwidth PCIe interconnect, CERAR reduces raw All-Reduce time by up to 47\%. For training a 1.4-billion-parameter Pythia model, it reduces training time by about 17\% for the same number of iterations while achieving essentially identical validation loss. On GPU clusters with high-bandwidth NVLink interconnects, CERAR performs worse in the All-Reduce phase and slightly worse in overall training time. However, the gap narrows for larger model sizes, suggesting that further gains may be possible with more optimized CERAR implementations.
Comments21 pages, 1 figure