发表机构
NVIDIA(英伟达)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一个端到端系统,通过扩展上下文并行和流水线并行支持大规模长上下文在线草稿协同训练,在高达122B模型规模上实现显著加速。
AI 中文摘要
投机解码加速了rollout生成,而rollout生成主导了强化学习(RL)后训练的成本。在线协同训练可以进一步提高草稿的准确性,从而带来更大的加速效果。然而,将这种方法扩展到具有长上下文的大模型协同训练面临两个障碍:(1)标准因果上下文并行(CP)实现不支持分支注意力,(2)目标特征跨越流水线并行(PP)阶段。我们通过一个用于大规模在线草稿协同训练的端到端系统解决了这两个问题。对于CP,我们通过将秩局部分支注意力与因果主序列注意力合并,扩展了打包的、负载均衡的之字形环注意力。对于PP,TapChannel通过独立路径跨阶段传输中间目标特征,不影响流水线调度。实验表明,协同训练的草稿紧密跟踪策略基线,同时在高达122B的模型规模上提供了显著的rollout和端到端加速。我们的CP设计在256K令牌下实现了强扩展,并相比先前工作节省了大量内存,而我们的PP传输开销适中。代码可在https://this https URL找到。
英文摘要
Speculative decoding accelerates rollout generation, which dominates the cost of reinforcement learning (RL) post-training. Online co-training can further increase the draft's accuracy, yielding greater speedups. However, scaling this approach to co-training on large models with long contexts poses two obstacles: (1) branch attention is unsupported by standard causal context-parallel (CP) implementations, and (2) target features span across pipeline-parallel (PP) stages. We address both with an end-to-end system for large-scale online draft co-training. For CP, we extend packed, load-balanced zigzag ring attention by merging rank-local branch attention with causal main-sequence attention. For PP, TapChannel transports intermediate target features across stages via a separate path, leaving the pipeline schedule unaffected. Experiments demonstrate that co-trained drafts closely track the policy baseline while delivering substantial rollout and end-to-end speedups across model scales up to 122B. Our CP design achieves strong scaling at 256K tokens with significant memory savings over prior work, and our PP transport incurs modest overhead. Code can be found at https://github.com/NVIDIA-NeMo/RL/issues/3698.
CommentsTechnical Report