arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

验证工具的覆盖范围决定其价值:AI编码智能体中验证面、产物质量与成本的对照研究

The reach of a verification tool decides its value: A controlled study of verification surface, artifact quality, and cost in AI coding agents

Achint Mehta

arXiv 2608.28795首次发表:更新:

AI 中文总结

该研究通过构建仅工具列表为受控变量的极简编码智能体,开展对照实验,发现验证工具的覆盖范围需匹配应用实际故障方式才能提升产物质量,且不同验证工具的成本效益存在差异。

AI 中文摘要

现代人工智能编码智能体可配备检查自身工作的工具,例如代码检查器(linter)、启动探针(boot probe)、shell、截图工具,我们将这组工具称为智能体的验证面。本研究探究在其他条件固定的情况下,仅扩大该验证面是否会使智能体产出软件的质量出现相应增长。我们构建了一个极简编码智能体,其工具列表为唯一受控变量,并用它在六个模型、八种工具配置下实现了1116个Web应用。盲态人类根据固定评分标准对每个应用进行评分,自动探针则对API可观测行为进行压力测试。验证的最低成本效益最先显现,即确保应用能正常启动:无工具时,约七分之一的构建完全无法启动;单个启动探针可消除几乎所有此类故障,其token成本约为完整shell的35%,而完整shell则将无工具时的成本乘以2.35。截图在错误可见的场景(如元素布局、交互)中帮助最大,但即便在此场景下,其相较于shell的增益也较为有限,且在对多次统计比较进行校正后不再显著。在故障仅可测量而非可见的场景(如在10万行列表上保持滚动流畅)中,截图无任何增益。验证工具仅在其覆盖应用实际故障方式的场景中才能提升输出产物的质量。

英文摘要

Modern artificial-intelligence coding agents can be equipped with tools for checking their own work e.g. a linter, a boot probe, a shell, a screenshot tool. We call this set the agent's verification surface. This study asks whether increasing only that surface, with everything else held fixed, produces a matching growth in the quality of the software the agent ships. We built a minimal coding agent whose tool list is the single controlled variable and used it to implement 1,116 web applications across six models and eight tool configurations. A condition-blind human graded every application against a frozen rubric, and automatic probes stress-tested the API-observable behaviors. Verification's cheapest benefit arrives first, which is to make sure that the application comes up. Without any tools, about one build in seven fails to launch at all and a single boot probe removes nearly all of these failures at roughly 35 percent of a full shell's token cost, while the full shell multiplies the no-tools cost by 2.35. Screenshots help most where mistakes are visible (e.g. element placement, interaction), though even there the gain over a shell is modest and does not survive correction for multiple statistical comparisons. In cases where failures can only be measured rather than seen, such as keeping scrolling smooth over a 100,000-row list, screenshots add nothing. A verification tool improves the output artifact only where its reach covers the way the application actually fails.

Comments27 pages, 13 figures, 10 tables. Submitted to IEEE Access. Data and code: https://doi.org/10.5281/zenodo.21961590

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑