发表机构
Alibaba Group; University of Science and Technology of China; Tsinghua University(阿里巴巴集团; 中国科学技术大学; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
QwenGyre 是一个端到端弹性强化学习框架,通过动态 GPU 分配和轨迹去重,高效训练超长时域智能体,在 NL2RepoBench 上提升 6.0%,并实现最高 1.85 倍加速。
AI 中文摘要
大型语言模型(LLM)智能体越来越多地承担超长时域(xlong)任务,其中单次执行可跨越数小时、数百次模型与环境交互,且每次 rollout 接近 100 万 token。将在线强化学习(RL)应用于此类执行面临两个基本挑战:(1)严重的执行方差和长时间的 rollout 延迟导致大量 GPU 空闲;(2)复杂的非线性分支产生大量轨迹冗余,严重削弱训练效率。为解决这些问题,我们提出了 QwenGyre,一个用于超长时域在线 RL 的端到端框架。QwenGyre 在不中断实时执行的情况下,在 rollout 和训练之间弹性重新分配 GPU,同时其轨迹处理器重构分支历史、对部分进展进行评分,并对冗余路径进行去重,以限制训练成本。在扩展到我们的旗舰模型 Qwen~3.8 2.4T(每次 rollout 包含 70 万 token)后,QwenGyre 在 48 步内在 NL2RepoBench 上取得了 6.0% 的绝对提升(52.5% 提升至 58.5%)。在我们对多样化训练数据集领域的评估中,QwenGyre 相比 Colocate 和 Async 分别实现了高达 1.85 倍和 1.78 倍的加速。
英文摘要
Large language model (LLM) agents increasingly undertake extreme-long (xlong) horizon tasks, where a single execution can span hours, hundreds of model--environment interactions, and nearly 1M tokens per rollout. Applying online reinforcement learning (RL) to such executions poses two fundamental challenges: (1) severe execution variance and prolonged rollout delays cause massive GPU idling; and (2) complex non-linear branching generates massive trajectory redundancy, crippling training efficiency. To address these, we presents QwenGyre, an end-to-end framework for xlong-horizon online RL. QwenGyre elastically reallocates GPUs between rollout and training without interrupting live executions, while its trajectory processor reconstructs branching histories, scores partial progress, and deduplicates redundant paths to bound training costs. Scaled to our flagship model, Qwen~3.8 2.4T, with 700K tokens per rollout, QwenGyre yields a 6.0% absolute gain on NL2RepoBench (52.5% $\to$ 58.5%) in 48 steps. Across our evaluations on diverse domains of training datasets, QwenGyre delivers up to $1.85\times$ and $1.78\times$ speedups over Colocate and Async, respectively.