基于补丁推理的软件智能体可扩展监督
Scalable Supervision for Software Agents via Patch Reasoning
- School of Data Science, The Chinese University of Hong Kong, Shenzhen(数据科学学院,香港中文大学(深圳))
- ByteDance(字节跳动)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对现有基于测试的软件智能体监督可扩展性不足的问题,提出R4P推理方法,在SWE-bench补丁验证中准确率达72.2%,训练的Mini-SE在Pass@1上较Qwen3-32B提升10.0%至26.2%。
AI中文摘要:
尽管语言模型智能体在软件工程领域取得了进展,但现有的基于测试的监督方式限制了其在实际问题上的可扩展性,原因有二:一是实际场景中高覆盖率测试天然稀少,二是构建和运行测试沙箱成本高且不稳定。为实现监督的可扩展,我们提出R4P,一种基于推理的方法,可提供与脚手架无关的奖励。R4P采用分组训练目标,使其能够相互验证多个补丁的修改情况,获得密集奖励以监督智能体,无需执行测试或依赖特定智能体轨迹。R4P在验证SWE-bench补丁时准确率达72.2%,与专有模型表现相当。为展示R4P的下游实用价值,我们设计并训练了无执行脚手架Mini-SE,通过R4P采用纯强化学习实现,Mini-SE的Pass@1达26.2%,较原始Qwen3-32B提升10.0%,在补丁选择的测试时扩展中,借助R4P可进一步提升至32.8%。稳定的扩展曲线表明,尽管存在不足,R4P仍能可靠地支持下游任务的大规模应用。
英文摘要:
While language model agents have advanced software engineering, existing test-based supervision is limiting its scalability on real-world issues. The reason is twofold: (1) high-coverage tests are naturally rare in the wild, and (2) building and running test sandbox is heavy and fragile. To unlock supervision scaling, we propose R4P, a reasoning-based method that provides scaffold-agnostic rewards. R4P uses a group-wise training objective, enabling it to verify multiple patches against each other's modification and gain a dense reward for supervising agents without executing tests or relying on specific agent trajectories. R4P achieves 72.2% Acc. for verifying patches from SWE-bench, competitive with proprietary models. To show the downstream practical utility of R4P, we design and train an execution-free scaffold, Mini-SE, with pure RL via R4P. Mini-SE achieves 26.2% Pass@1, showing a 10.0% improvement over the original Qwen3-32B, and can be further improved to 32.8% with R4P for test-time scaling on patch selection. The stable scaling curves illustrate that though imperfect, R4P can still reliably support downstream tasks at scale.