arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

将AI智能体建立在契约基础上:对规范驱动测试生成的实证评估

Grounding AI Agents in Contracts: An Empirical Evaluation of Spec-Driven Test Generation

Michele Tufano, James McClure, José Cambronero, Runxiang Cheng, Sherry Y. Shi, Renyao Wei, Dorothy Chen, Franjo Ivančić, Livio Dalloro, Pat Rondon

arXiv 2608.17177首次发表:更新:

AI 中文总结

本研究针对LLM智能体生成测试时易遗漏契约相关边界的问题,提出规范驱动测试生成方法,经Google生产缺陷评估,其缺陷检测率、分支覆盖率均优于基线,测试套件质量也显著更高。

AI 中文摘要

基于大语言模型(LLM)的智能体越来越多地用于编码任务,它们的表现已优于许多经典方法,并扩展到了仓库级任务,例如测试生成。然而,当直接提示这些智能体生成测试时,它们可能无法推理代码及其底层契约,从而遗漏影响测试质量的边界情况和行为界限。为解决这一局限,我们提出了规范驱动测试生成方法,即指导智能体首先推理并明确记录代码的前置条件、后置条件和未定义行为。这种中间半形式化规范充当认知支架,用于指导后续的测试生成。我们对来自Google的生产缺陷进行评估,结果显示,与传统测试生成智能体基线相比,该规范驱动智能体的缺陷检测率提升了9.8个百分点(p=0.0352),分支覆盖率提升了2.5个百分点(p=0.0034)。我们进一步使用“大语言模型作为评判者”的方式表明,规范驱动智能体生成的测试套件在77.8%的案例中优于基线生成的测试,在56.7%的案例中优于人工编写的测试,且在遵循最佳实践、可读性和边界情况覆盖方面均有提升。

英文摘要

LLM-based agents are increasingly used for coding tasks, where they have outperformed many classical approaches and scaled to repository-level tasks, such as test generation. However, when directly prompted to generate tests, these agents can fail to reason about the code and its underlying contracts, thereby missing edge cases and behavioral boundaries that affect test quality. To address this limitation, we propose Spec-Driven Test Generation, where we instruct an agent to first reason about -- and explicitly document -- code pre-conditions, post-conditions, and undefined behaviors. This intermediate semi-formal specification acts as a cognitive scaffold to guide subsequent test generation. Our evaluation on production bugs from Google shows that the spec-driven agent can deliver a 9.8 percentage points ($p = 0.0352$) improvement in bug detection rate and a 2.5 percentage point ($p = 0.0034$) improvement in branch coverage, compared to a traditional test generation agent baseline. Using LLM-as-a-Judge, we further show that test suites generated by the spec-driven agent are superior to the baseline and human-authored tests in 77.8% and 56.7% of the cases, respectively, and demonstrated improvements on following best practices, readability, and edge-case coverage.

CommentsTo appear in Proceedings of the 1st International Workshop on Specification-Driven Development Life Cycle (SpecOps 2026), co-located with SPLASH 2026

DOI:10.1145/3842652.3843195

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑