发表机构
Georgia Institute of Technology; University of Colorado Boulder(佐治亚理工学院; 科罗拉多大学博尔德分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究以BSG-VA方法分析LLM修复智能体的验证证据,发现近半数阳性测试无缺陷区分信息,缺陷对比反馈可减少证据不足的部署,其效果幅度仍需进一步验证。
AI 中文摘要
当修复智能体运行测试并看到测试通过时,该结果会被视为关于所报告缺陷的证据。本研究衡量这种处理方式的合理性程度。BSG-VA(缺陷状态/候选状态/黄金修复验证分析)在精确的工作树状态下捕获每个验证命令,提取仅用于测试的补丁,并在原始缺陷代码(B)、候选状态(S)和开发者黄金修复(G)上重放该命令。捕获的结果和重放结果为每个事件分配一个证据角色,从与黄金修复对齐的缺陷区分型,到仅回归型,再到误导型。在110个任务的643次部署中的3730个事件里,46.0%的阳性可比事件不携带缺陷区分信息;23.8%的无反馈注入的基准部署,最终补丁的全部阳性证据基础都属于此类。一项三臂实验测试将B重放结果返回给智能体是否会改变这一模式。与注意力匹配的提醒相比,缺陷对比反馈使证据不足的部署减少了7.8个百分点(p=0.0029),并使缺陷区分证据增加了7.4个百分点(p=0.011),且未对修复成功率造成可检测的损失。这两个估计值均低于预先设定的10个百分点的最小感兴趣效应量,因此实际幅度仍不确定。约三分之一的改进源于提醒本身;在两个探索性复现中,通过改变支架和模型,仅在无约束工具使用循环下使用gpt-5.6-sol时,B重放内容才会带来可检测的增量。BSG-VA事后适用于任何保留所需代码状态和执行环境的可重放修复轨迹。
英文摘要
When a repair agent runs a test and sees it pass, the result is treated as evidence about the reported defect. We measure how often that treatment is warranted. BSG-VA (buggy-state/candidate-state/gold-fix validation analysis) captures each validation command at its exact working-tree state, extracts a test-only patch, and replays the command on the original buggy code (B), the candidate state (S), and the developer gold fix (G). The captured outcome and the replay results assign every event an evidence role, from gold-aligned bug-discriminating through regression-only to misleading. Across 3,730 events in 643 rollouts on 110 tasks, 46.0% of positive comparable events carry no bug-discriminating information; 23.8% of baseline rollouts, with no feedback injected, close with a patch whose entire positive evidence base is of this kind. A three-arm experiment tests whether returning the B-replay outcome to the agent changes this pattern. Bug-contrast feedback reduces evidence-inadequate closure by 7.8 percentage points relative to an attention-matched reminder (p = 0.0029) and raises bug-discriminating evidence by 7.4 points (p = 0.011), with no detectable cost to repair success. Both estimates fall below the prespecified 10-percentage-point smallest effect size of interest, so practical magnitude remains uncertain. Roughly a third of the improvement traces to the reminder alone; across two exploratory replications, varying the scaffold and the model, the B-replay content adds a detectable increment only with gpt-5.6-sol under the unconstrained tool-use loop. BSG-VA applies post hoc to any replayable repair trajectory that preserves the required code states and execution environment. Keywords: program repair agents, validation evidence, test adequacy, large language models, software quality, controlled experiment.