发表机构
University of California, Berkeley; Virginia Tech(加州大学伯克利分校; 弗吉尼亚理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究探究弱评审者能否审计强编码智能体,通过执行证据和级联测试提升缺陷捕获与降低过度拒绝,发现评审者规模非关键,生成可靠检查是主要瓶颈。
AI 中文摘要
编码智能体可能返回看似合理但遗漏了所需行为的补丁。这些失败难以审查,因为冗长的追踪记录和自信的摘要常常掩盖了遗漏之处。我们探究名义上较弱的评审者何时能可靠地判断补丁是否解决了问题。我们研究了来自三个智能体的411条带执行标签的追踪记录和101个受控案例。在154条GPT-5.4追踪记录上,结构化但未经检查的证据同时提高了缺陷捕获率和过度拒绝率。随后我们提供官方执行证据作为上限诊断。在按评审者为两种格式各选择并固定一种后,六位评审者中有五位在122条保留追踪记录上同时提高了两项指标;两位对每条追踪记录都分类正确。评审者规模并非质量的一致预测因素。由于部署中无法获得官方测试,我们还评估了一个冻结的级联流程,该流程结合了补丁引起的静态错误和首先生成在未打补丁仓库上失败的测试。在121条评分的保留GPT-5.4追踪记录和59条Gemini追踪记录上,其覆盖率分别为0.89和0.86,风险为0.33和0.26,捕获率为0.76和0.80,过度拒绝率为0.66和0.67。大多数错误拒绝发生在未解决案例到达评审者时。官方执行证据表明,在存在决定性检查时,弱评审具有潜力。在没有官方测试的情况下生成同样可靠的检查仍是主要瓶颈。
英文摘要
Coding agents can return plausible patches that omit required behavior. These failures are hard to review because long traces and confident summaries often hide what was missed. We ask when a nominally weaker reviewer can reliably decide whether a patch solves its issue. We study 411 execution-labeled traces from three agents and 101 controlled cases. On 154 GPT-5.4 traces, structured but unchecked evidence raises both defect catch and over-rejection. We then provide official execution evidence as an upper-bound diagnostic. After choosing and freezing one of two formats per reviewer, five of six reviewers improve both rates on 122 held-out traces; two classify every trace correctly. Reviewer size is not a consistent predictor of quality. Because official tests are unavailable in deployment, we also evaluate a frozen cascade with patch-caused static errors and generated tests that first fail on the unpatched repository. On 121 scored held-out GPT-5.4 traces and 59 Gemini traces, its coverage is 0.89 and 0.86, risk is 0.33 and 0.26, catch is 0.76 and 0.80, and over-rejection is 0.66 and 0.67. Most false rejections occur when unresolved cases reach the reviewer. Official execution evidence shows the potential of weak review when decisive checks are available. Producing equally reliable checks without official tests remains the main bottleneck.