发表机构
National University of Singapore(新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对可变长序列深度学习训练的效率与易用性权衡难题,提出以数据为中心的并行(DCP)方法,通过动态调整运行时设置实现2.88倍加速,仅需10行代码即可集成到任意模型中。
AI 中文摘要
在可变长序列上训练深度学习模型会带来显著的计算挑战。现有方法需在效率与易用性之间进行艰难权衡:简单方法采用静态配置,会导致工作负载不平衡、效率低下;复杂方法则会引入大量复杂性,且需为新模型进行代码修改。为打破这一权衡,我们提出以数据为中心的并行(Data-Centric Parallel,DCP)方法,其核心原则是让数据本身驱动运行时,通过根据每个批次的序列长度动态调整直接运行时设置(如并行规模、梯度累积、重计算)来实现。实验结果表明,我们的方法在32块H200 GPU上实现了最高2.88倍的加速。该方法专为泛化性设计,仅需10行代码即可集成到任何模型中。我们预计这种简单却有效的方法将成为稳健的基准,并推动可变长序列分布式训练的未来进展。
英文摘要
Training deep learning models on variable long sequences poses significant computational challenges. Existing methods force a difficult trade-off between efficiency and ease-of-use. Simple approaches use static configurations that cause workload imbalance low efficiency, while complex methods introduces significant complexity and code change for new models. To break this trade-off, we introduce Data-Centric Parallel (DCP). Its core principle is to let the data itself drive the runtime. It achieves this by dynamically adjusting direct runtime settings (e.g., parallel size, gradient accumulation, recomputation) based on each batch's sequence length. Empirical results demonstrate that our method achieves up to a 2.88$\times$ speedup on 32 H200 GPUs. Designed for generalization, it can be integrated into any model with 10 lines of code. We anticipate this simple yet effective approach will serve as a robust baseline and facilitate future advancements in distributed training for variable long sequences.