发表机构
Department of Computer Science Florida International University Miami, FL, USA; Queens College; The Graduate Center City University of New York New York, NY, USA(计算机科学系佛罗里达国际大学; 皇后学院; 纽约市立大学研究生中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大规模机器学习分布式训练中掉队工作节点致训练慢的问题,提出流水线梯度编码方法,为分数重复和循环重复开发流水线版本,经实验验证该方法能显著减少训练时间并加速收敛。
AI 中文摘要
在大规模机器学习中,分布式训练通常涉及多个工作节点在不同数据集分区上评估模型梯度。一个常见挑战是存在掉队工作节点,这可能显著减缓训练速度。传统梯度编码(GC)通过在工作节点间复制数据集分区来解决此问题,允许替换掉队节点缺失的梯度。然而,GC要求工作节点在每一步对多个数据集分区评估梯度,可能增加总体训练时间。本文提出流水线GC,使梯度评估跨多步分段进行,每个工作节点每步仅在单个数据集分区上评估梯度。我们为GC中的两种代表性数据集放置方案——分数重复(FR)和循环重复(CR)开发了流水线版本,并证明了两者的收敛保证。通过在云基础设施上的广泛模拟和实验,我们的方案与GC及其他基线相比,不仅显著减少训练时间,还加速了收敛。
英文摘要
In large-scale machine learning, distributed training commonly involves multiple workers evaluating the gradients of the model on different dataset partitions. A common challenge is the presence of straggling workers, which may significantly slow down training. Traditional gradient coding (GC) addresses this by duplicating dataset partitions across workers, allowing for the replacement of missing gradients from stragglers. However, GC requires workers to evaluate gradients on multiple dataset partitions in each step, potentially increasing overall training time. In this paper, we propose to pipeline GC, such that gradient evaluation is segmented across multiple steps and each worker evaluates gradients on just a single dataset partition per step. We develop the pipelined version for fractional repetition (FR) and cyclic repetition (CR), two representative dataset placement schemes in GC, and prove convergence guarantees for both. Through extensive simulations and experiments on cloud infrastructure, our schemes not only significantly reduce training time but also accelerate convergence compared to GC and other baselines.
CommentsThis is the extended version of the paper accepted at the IEEE Information Theory Workshop (ITW) 2026