NSP:通过嵌套序列并行加速变长LLM训练
NSP: Accelerating Variable-Length LLM Training via Nested Sequence Parallelism
浏览论文内容
中文总结 AI 辅助
NSP通过嵌套不同大小的序列并行组,在共享GPU上平衡长尾变长序列训练的通信与负载,无需修改模型,显著提升端到端吞吐量。
中文摘要 AI 辅助
在长尾语料上进行长上下文LLM训练面临通信与平衡之间的核心权衡。这种序列长度异质性使得任何单一的序列并行(SP)度都无法很好地适应该工作负载:较小的度会导致少数长序列严重不平衡,而较大的度则迫使占主导地位的众多短序列承担过多的通信开销。现有的动态SP系统在一个批次内混合不同的SP度,但为了同时运行多个组,它们将GPU划分为不相交的组,这重新引入了组间的不平衡,并迫使采用昂贵的微批次变通方法。我们提出了NSP,一个序列并行训练系统,通过在单个训练迭代内在共享GPU上嵌套不同大小的SP组来解决这一权衡。这使得长序列可以使用较大的SP组,而短序列则保持在较小的组上,因此通信仅在需要的地方产生,并且负载按GPU而非按组进行平衡。NSP通过一个树状结构的路由规划器在内存约束下分配序列,以及一个通过层间阶段流和树级重计算利用所得层次结构的执行器来实现这一想法。NSP支持常见的SP后端,且无需修改模型。我们在内部生产GPU集群上,针对多个长尾数据集上的Qwen3-MoE工作负载(上下文长度高达384K token)评估了NSP。在这些设置中,NSP持续提升了端到端训练吞吐量,比静态SP最高提升1.48倍,比FlexSP最高提升1.16倍。
英文摘要
Long-context LLM training on long-tailed corpora faces a central communication--balance tradeoff. Such sequence-length heterogeneity makes any single sequence-parallelism (SP) degree a poor fit for the workload: a small degree leaves the few long sequences badly imbalanced, while a large degree forces the many short sequences that dominate the workload to pay excessive communication. Existing dynamic-SP systems mix SP degrees within a batch, but to run several groups at once they partition the GPUs into disjoint groups, which reintroduces imbalance across groups and forces costly micro-batch workarounds. We present NSP, a sequence-parallel training system that resolves this tradeoff by nesting differently sized SP groups on shared GPUs within a single training iteration. This lets long sequences use larger SP groups while keeping short sequences on smaller ones, so communication is incurred only where needed and load is balanced per GPU rather than per group. NSP realizes this idea with a tree-structured routing planner that assigns sequences under memory constraints and an executor that exploits the resulting hierarchy through inter-level phase streaming and tree-level recomputation. NSP supports common SP backends and requires no model changes. We evaluate NSP on Qwen3-MoE workloads with up to 384K-token contexts across multiple long-tail datasets on an internal production GPU cluster. Across these settings, NSP consistently improves end-to-end training throughput, outperforming Static SP by up to 1.48x and FlexSP by up to 1.16x.
发表机构
- ByteDance Seed(字节跳动Seed)
机构由 AI 辅助整理,请以论文原文为准。