arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于双向重构-验证框架的代码智能体独立补丁验证

Independent Patch Verification for Coding Agents with a Bidirectional Reconstruct-and-Verify Framework

Chenglin Li, Yisen Xu, Zehao Wang, Shin Hwei Tan, Tse-Hsun, Chen

arXiv 2608.08950首次发表:更新:

AI 中文总结

本研究针对代码智能体生成补丁后缺乏独立验证的问题,提出无需训练的RETRACE双向重构-验证框架,在SWE-bench Verified等基准上显著提升了代码补丁的正确性。

AI 中文摘要

由大语言模型驱动的自主代码智能体现在可以直接从bug报告生成代码补丁,但存在一个根本缺陷:补丁生成后,没有机制独立验证其是否真正解决了报告的问题。现有研究试图通过迭代自精化和推理时缩放来解决该问题,但这些方法要么在生成补丁的同一视角下审查补丁,要么扩大候选补丁生成范围却不验证单个补丁,均未提供评估补丁正确性的显式验证信号。我们提出RETRACE,一种无需训练的补丁生成后验证框架,通过双向重构与协调推导该信号。当代码智能体为某问题生成候选补丁时,RETRACE执行正向重构,从问题和智能体的轨迹构建显式修复依据;随后反向重构仅从补丁及其轨迹(不访问原始问题)独立推断该补丁看似解决的问题描述,并将此重构结果与原始问题对比以生成对齐裁决;协调阶段则检查正向依据与补丁的一致性,诊断任何不对齐的来源,要么提交补丁,要么生成针对性的修订指导。在SWE-bench Verified上使用两个主干模型(GPT-5-mini和MiniMax-2.5)评估,RETRACE在mini-SWE-agent支架上分别将Pass@1提升7.0%和3.6%,在OpenHands上无需修改即可实现相当的提升。消融实验显示,正向和反向阶段均对整体提升有贡献,添加协调阶段可进一步提升效果。

英文摘要

Autonomous coding agents powered by large language models can now generate code patches directly from bug reports, but a fundamental gap remains: once a patch is produced, no mechanism independently verifies whether it truly resolves the reported problem. Prior work has sought to address this through iterative self-refinement and inference-time scaling, but these approaches either review the patch under the same interpretation that produced it or broaden candidate generation without verifying individual patches, and neither provides an explicit verification signal for assessing patch correctness. We propose RETRACE, a training-free post-generation verification framework that derives such a signal through bidirectional reconstruction and reconciliation. When a coding agent generates a candidate patch for an issue, RETRACE performs forward reconstruction to build an explicit repair rationale from the issue and the agent's trajectory; backward reconstruction then independently infers, from the patch and its trajectory alone and without access to the original issue, a description of the problem the patch appears to address, and compares this reconstruction against the original issue to produce an alignment verdict; a reconciliation stage then checks the consistency between the forward rationale and the patch, diagnoses the source of any misalignment, and either submits the patch or produces targeted revision guidance. Evaluated on SWE-bench Verified with two backbones (GPT-5-mini and MiniMax-2.5), RETRACE lifts Pass@1 by 7.0% and 3.6% respectively on the mini-SWE-agent scaffold, and delivers comparable gains on OpenHands without modification. Ablation experiments show that both the forward and backward stages contribute to the overall improvement and that adding reconciliation yields further gains.

Comments8 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑