AI 中文总结
提出反事实回放(CRR),利用可分支环境生成步骤级回报对比,无需过程标签或奖励模型,显著提升SWE智能体在多个基准上的pass@1。
AI 中文摘要
仅基于结果的强化学习为软件工程(SWE)智能体提供终局成功信号,但对中间决策的直接指导很少。我们提出反事实回放(CRR),一种训练时程序,利用可分支的可执行环境来获得步骤级别的回报对比。CRR选择少量决策点,恢复每个状态,采样替代动作,并在策略下将分支向前推进。它保留实际训练轨迹,并将选定步骤的优势替换为其终局回报与采样反事实回报之间的差值。该方法不需要人工过程标签或学习的过程奖励模型;免费指的是这些监督成本,而非回放计算。使用14B策略,CRR在SWE-bench Verified、SWE-bench Live和SWE-rebench上提高了pass@1,并与过程奖励和轨迹搜索方法相结合。在SWE-bench Verified上,相同硬件上的等墙钟时间比较显示,CRR达到41.7%,而扩展的仅结果GRPO为36.7%,在包含分支开销的情况下提升了5.0个百分点。这些结果适用于具有可靠、低成本状态恢复的环境;随机延续以及昂贵或不完美的回放仍是局限性。
英文摘要
Outcome-only reinforcement learning gives software engineering (SWE) agents a terminal success signal but little direct guidance about intermediate decisions. We introduce Counterfactual Rollout Replay (CRR), a training-time procedure that uses forkable executable environments to obtain step-level return contrasts. CRR selects a small set of decision points, restores each state, samples an alternative action, and rolls the branch forward under the policy. It retains the realised training trajectory and replaces the advantage at selected steps with the difference between its terminal return and the sampled counterfactual return. The method needs no human process labels or learned process reward model; free refers to those supervision costs, not replay compute. With a 14B policy, CRR improves pass@1 on SWE-bench Verified, SWE-bench Live, and SWE-rebench, and combines with process-reward and trajectory-search methods. On SWE-bench Verified, an equal-wall-clock comparison on the same hardware yields 41.7% versus 36.7% for extended outcome-only GRPO, a 5.0-point gain with fork overhead included. These results apply to environments with affordable, reliable state restoration; stochastic continuations and expensive or imperfect replay remain limitations.
CommentsAccepted at NeurIPS 2026. Includes additional experiments and analysis