发表机构
NVIDIA; Tsinghua University(英伟达; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FC-SWE提出失败条件强化学习框架,利用失败补丁和验证器反馈作为上下文训练恢复轨迹,通过轨迹局部奖励和活跃集优势估计改进GRPO,在SWE-bench Verified上显著提升解决率。
AI 中文摘要
仓库级软件工程(SWE)是一个具有挑战性的长时程任务:智能体必须在扩展交互中进行推理、使用工具,并适应有状态的环境。近期工作使用诸如组相对策略优化(GRPO)等强化学习方法训练SWE智能体,该方法针对每个问题独立采样多条轨迹,测试生成的补丁,并在固定组内比较最终奖励。然而,这种训练设置并未将来自失败补丁的验证器反馈作为后续尝试的上下文重用,尽管该反馈包含有关出错原因的宝贵诊断信息。在恢复轨迹上进行训练具有挑战性,因为前一个结果决定了是否生成下一个轨迹,而失败的执行决定了其条件上下文。我们提出FC-SWE,一种失败条件强化学习框架,将恢复尝试纳入策略训练。当补丁未通过验证时,FC-SWE将仓库恢复到其原始任务状态,并使用失败的补丁和验证器反馈作为恢复轨迹的上下文。FC-SWE通过两种机制将GRPO适配到这些完整的、多轮工具使用轨迹链中。轨迹局部奖励保留每次尝试的验证器结果,防止恢复成功奖励先前失败的补丁。活跃集优势估计从针对同一问题实际执行的所有初始和恢复轨迹中形成比较组,因此失败的尝试保留在组中,而未执行的尝试被排除。在验证器辅助协议下,在SWE-bench Verified的全部500个任务上,使用Qwen3.5-4B和SWE-agent的FC-SWE实现了41.7%的Resolved@1和52.8%的Resolved@2,而GRPO分别为38.9%和48.5%。尽管每个链最多训练两次尝试,FC-SWE在十一次尝试的测试时预算下达到了70.7%的Resolved@11。
英文摘要
Repository-level software engineering (SWE) is a challenging long-horizon setting: agents must reason over extended interactions, use tools, and adapt to stateful environments. Recent work trains SWE agents with reinforcement learning methods such as Group Relative Policy Optimization (GRPO), which independently sample multiple trajectories per issue, test the resulting patches, and compare terminal rewards within a fixed group. However, this training setup does not reuse verifier feedback from failed patches as context for subsequent attempts, even though this feedback contains valuable diagnostic information about what went wrong. Training on recovery trajectories is challenging because the preceding outcome determines whether the next trajectory is generated, while the failed execution determines its conditioning context. We introduce FC-SWE, a failure-conditioned RL framework that incorporates recovery attempts into policy training. After a patch fails verification, FC-SWE restores the repository to its original task state and uses the failed patch and verifier feedback as context for a recovery trajectory. FC-SWE adapts GRPO to these chains of complete, multi-turn tool-use trajectories through two mechanisms. Trajectory-local rewards preserve each attempt's verifier outcome, preventing recovery success from rewarding an earlier failed patch. Active-set advantage estimation forms a comparison group from all initial and recovery trajectories actually executed for the same issue, so failed attempts remain in the group while unexecuted attempts are excluded. On all 500 SWE-bench Verified tasks under a verifier-assisted protocol, FC-SWE with Qwen3.5-4B and SWE-agent achieves 41.7% Resolved@1 and 52.8% Resolved@2, compared with 38.9% and 48.5% for GRPO. Although trained with at most two attempts per chain, FC-SWE reaches 70.7% Resolved@11 under an eleven-attempt test-time budget.
Comments23 pages, 8 figures, 6 tables