发表机构
Stanford University; Yonsei University; Korea University; Bespoke Labs(斯坦福大学; 延世大学; 高丽大学; Bespoke Labs)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ThunderSyncRL通过梯度流实现智能体强化学习的无损加速,消除同步训练中的空闲气泡,同时避免异步训练的策略陈旧性,在SWE-bench Verified和Terminal Bench 4.0上实现最高1.9倍加速和2.47个百分点的性能提升。
AI 中文摘要
语言模型正从生成答案转向在交互式环境中追求长期目标。对这些智能体进行后训练需要长且异构的轨迹,而同步系统会让学习引擎在滚动和验证完成之前处于空闲状态。为消除这些流水线气泡,异步训练在更新之间重叠滚动和学习,但代价是策略陈旧性。我们提出ThunderSyncRL,它在所有必需输入确定后立即开始梯度计算,且不引入策略陈旧性。对于组相对策略优化(GRPO),ThunderSyncRL在每条轨迹的奖励到达时立即计算其得分梯度,而无需等待整个组。对于在策略蒸馏(OPD),它在工具调用在沙箱中运行时,为每个完成的智能体回合的教师评分动作计算梯度。我们证明梯度流产生的GRPO和OPD更新与批量同步训练相同,且不改变任一目标。在SWE-bench Verified和Terminal Bench 4.0上,我们将模型训练到相同性能,速度比同步训练快达1.9倍。在零策略陈旧性下,ThunderSyncRL在固定预算下也比异步训练性能高出最多2.47个百分点。
英文摘要
Language models are moving beyond generating answers to pursuing long-horizon goals in interactive environments. Post-training these agents requires long, heterogeneous trajectories, and synchronous systems leave learner engines idle until rollout and verification finish. To squeeze out these pipeline bubbles, asynchronous training overlaps rollout and learning across updates, but comes at the cost of policy staleness. We introduce ThunderSyncRL, which starts gradient computation as soon as all required inputs are fixed, without policy staleness. For group relative policy optimization (GRPO), ThunderSyncRL computes each trajectory's score gradient as soon as the reward for that trajectory arrives, without waiting for the group. For on-policy distillation (OPD), it computes gradients for each completed agentic turn's teacher-scored actions while tool calls run in the sandbox. We prove that gradient streaming produces the same GRPO and OPD updates as batch-synchronous training, without changing either objective. On SWE-bench Verified and Terminal Bench 4.0, we train models to the same performance up to $1.9 \times$ faster than synchronous training. With zero policy staleness, ThunderSyncRL also outperforms asynchronous training at a fixed budget by up to $2.47$ percentage points.