超越未通过测试:协同生成的错误重现测试与修复的迭代强化
Beyond Fail-to-Pass: Iterative Hardening of Co-Generated Bug Reproduction Tests and Fixes
查看机构详情
- Nanjing University(南京大学)
- Peking University(北京大学)
- Microsoft(微软公司)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
研究自动化程序修复中错误重现测试问题,指出仅用未通过到通过标准不足。提出CoHarden框架,先生成测试再迭代强化测试与修复,实验证明该框架在解决率等方面优于现有基线。
中文摘要 AI 辅助
大型语言模型使自动化程序修复在处理现实世界错误时更实用,但直接从错误报告进行修复仍受限。错误重现测试将错误报告转化为可执行的特定错误信号以指导修复和验证候选补丁。现有工作将其生成作为自动化程序修复的核心子问题,主要用未通过到通过标准评估。我们发现仅该标准不足以改进下游修复,一些此类测试宽松,会接受似是而非的错误补丁。我们将其分为严格和宽松两类来形式化这一缺失的质量维度,发现只有前者能持续提高修复成功率。我们还发现协同生成会引入测试-修复错误耦合。基于此,我们提出CoHarden框架,先生成测试,然后迭代强化测试和修复。实验表明CoHarden在SWE-bench Verified上达到69.4%的已解决率和78.9%的未通过到通过率,优于最强的仅修复和协同生成基线。
英文摘要
Large language models (LLMs) have made automated program repair (APR) increasingly practical for real-world bugs, but repairing directly from bug reports remains underconstrained. Bug reproduction tests (BRTs) help close this gap by turning a bug report into an executable, bug-specific signal that can guide repair and validate candidate patches. Existing work has therefore studied BRT generation as a core subproblem in APR and mainly evaluates a generated BRT using the fail-to-pass (F->P) criterion, which requires the test to fail on the buggy code but pass on the golden fix. We show that F->P alone is insufficient when the goal of a BRT is to improve downstream repair. In particular, some F->P BRTs are lax, reproducing the observed symptom yet still admitting plausible-but-incorrect patches. We formalize this missing quality dimension by separating F->P BRTs into rigorous and lax ones, and show empirically that only the former consistently improve repair success. We further find that co-generation introduces test--fix error coupling, where the in-trajectory fail-to-pass (F->P) check can pass even when both the generated patch and generated test are wrong. Based on these findings, we propose CoHarden, a co-generation framework that uses the Lax signal as an in-loop convergence criterion. CoHarden first generates a test before any fix, then iteratively hardens the test and fix against surviving mutation patches until the generated test no longer admits Lax regressions. Experiments show that CoHarden reaches 69.4% Resolved and 78.9% F->P on SWE-bench Verified, outperforming the strongest fix-only and cogeneration baselines by +9.6 and +7.9 percentage points in Resolved, respectively, with consistent gains across LLM backbones and benchmarks.