arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LongStraw:在固定GPU预算下超越200万个令牌的长上下文强化学习

LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget

Changhai Zhou, Kieran Liu, Yuhua Zhou, Qian Qiao, Jun Gao, Harry Zhang, Irvine Lu, Nolan Ho, Lucian Li, Andrew Lei, Cleon Cheng, Steven Chiang, Yihang Zeng, Di Zhang, Rio Yang, Kaijie Chen, Andrew Chen, Pony Ma, Weizhong Zhang, Cheng Jin

arXiv 2607.14952首次发表:更新:

AI 中文总结

研究针对推理上下文长度与强化学习后训练的差距问题,提出LongStraw架构感知执行堆栈,用GRPO优化,在固定GPU预算下实现百万令牌强化学习后训练,通过实验验证了其执行能力。

AI 中文摘要

推理上下文长度与强化学习后训练之间的差距日益增大:推理系统接近百万令牌上下文,而后训练工作负载通常保持在256K令牌或以下,并在部署时依赖长度泛化。对于智能体来说,这一差距尤为重要。LongStraw是一种在固定GPU预算下用于百万令牌强化学习后训练的架构感知执行堆栈,通过组相对策略优化(GRPO)实例化。它在不使用自动求导的情况下评估共享提示,仅保留后续令牌所需的特定于模型的状态,并一次重放一个短响应分支,以增加重放时间为代价减少实时训练图。我们在混合循环和全注意力的Qwen3.6 - 27B以及压缩注意力专家混合模型GLM - 5.2上实现了它。在八个H20 GPU上,LongStraw在210万个位置完成了分组Qwen评分和响应反向传播,对于2组和8组的情况;增加组大小仅增加0.21GB的峰值分配内存,而单独的压力测试达到446万个位置。在32个H20 GPU上,我们验证了GLM - 5.2所有78层上210万个令牌提示的端到端LongStraw执行路径。这些实验确定了执行能力而非完整的训练正确性,因为捕获的提示状态是分离的,并且一些分布式前向和梯度合成路径仍然不完整。

英文摘要

Long-context RL post-training is constrained by the lifetime of state and gradients, not attention cost alone. In GRPO, one multi-million-token prompt must serve old-policy and reference scoring plus multiple policy responses, while conventional autograd keeps the prompt graph and all response graphs live alongside model weights, caches, and distributed communication buffers. We present LongStraw, an objective-aware, architecture-aware system for resident-state virtualization, response replay, and distributed-gradient execution. Its transaction captures the shared prompt without autograd, retains only the architecture-required state on explicitly owned pages, restores that state for each group member, scores old/reference branches without a graph, replays one policy response at a time with autograd, and accumulates the resulting gradients before one distributed finalization and optimizer step. This schedule bounds the live training graph by the response suffix while reusing the expensive prompt computation across the complete GRPO group. We instantiate this design for two incompatible model structures. Qwen3.6-27B combines 48 recurrent GDN layers with 16 full-attention layers; LongStraw keeps the compact recurrent state and physically CP8-sharded KV pages, composes global attention through cross-rank LSE/output merging, and performs blockwise response replay. GLM-5.2 combines a 78-layer MLA/DSA attention stack with a 256-expert, top-8 MoE tail. Its implementation keeps CP-sharded MLA latent pages and DSA indexer-key pages in CPU memory, stages one layer at a time, reconstructs IndexShare-aware global sparse selection over CP32, and dispatches routed response tokens over EP32. The two paths share one transaction contract while specializing the retained state, replay operator, and collective communication to the architecture...

Comments44 pages, 10 figures, 11 tables. Code: https://github.com/MindLab-Research/longstraw

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑