arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

智能拉取请求的测试覆盖分析

Test Coverage Analysis of Agentic Pull Requests

Atish Kumar Dipongkor, Talank Baral, Wing Lam, Kevin Moran

arXiv 2607.18057首次发表:更新:

AI 中文总结

研究人工智能编码代理生成的拉取请求的测试覆盖情况,分析4882个相关请求,发现代理包含测试更改比例低,现有测试覆盖不足,代理编写测试仅在少数请求中提升覆盖,错误处理结构测试严重不足,为此提出相关改进措施。

AI 中文摘要

人工智能编码代理越来越多地在极少人工干预的情况下提交完整的拉取请求,使软件开发从人工智能辅助工作流程转向自主工作流程。随着这些代理的日益普遍,确保它们生成的代码通过现有测试或代理编写的测试得到充分测试对于防止回归至关重要,但目前对智能拉取请求中的测试了解甚少。为填补这一空白,我们分析了来自AIDev数据集的4882个由五个编码代理生成的拉取请求(532个Java和4350个Python拉取请求)。我们研究了代理包含测试更改的频率,以及现有测试和代理编写的测试对代码更改的覆盖程度。代理仅在49.6%的更改测试文件中的拉取请求中包含测试更改。现有测试提供的安全网不完整:它们覆盖了Java中代理更改的可执行行的61.5%,在Python中仅覆盖了27.0%,其中64.8%的拉取请求没有任何现有测试执行的更改行。代理编写的测试提高了对现有测试的覆盖,但仅在少数拉取请求中:35.9%的Java和22.5%的Python代码+测试拉取请求显示覆盖增加。在两种语言中,错误处理结构(如try和catch块)始终是测试不足的,在Java中的未命中率达到86.0%,在Python中达到81.0%。这些发现促使人们采用覆盖感知开发实践、为编码代理建立覆盖反馈循环,以及衡量测试质量的评估基准,以更好地帮助代理可靠地测试自己的代码。

英文摘要

AI coding agents increasingly submit complete pull requests (PRs) with minimal human intervention, shifting software development from AI-assisted to autonomous workflows. As these agents become more prevalent, ensuring the code they generate is adequately tested, by existing tests or by tests the agents write, is critical to preventing regressions, yet little is known about testing in agentic PRs. To address this gap, we analyze 4882 agent-generated PRs from the AIDev dataset (532 Java and 4350 Python PRs) produced by five coding agents. We study (i) how often agents include test changes and (ii) how well covered are code changes by existing and agent-written tests. Agents include test changes in only 49.6% of PRs that change code under test files. Existing tests provide an incomplete safety net: they cover 61.5% of agents' changed executable lines in Java and only 27.0% in Python, where 64.8% of PRs have no changed line executed by any existing test. Agent-written tests improve coverage over existing tests, but only in a minority of PRs: 35.9% of Java and 22.5% of Python Code + Tests PRs show a coverage gain. Across both languages, error-handling constructs (e.g., try and catch blocks) are the most consistently under-tested, with miss rates reaching 86.0% in Java and 81.0% in Python. These findings motivate coverage-aware development practices, coverage feedback loops for coding agents, and evaluation benchmarks that measure test quality to better help agents reliably test their own code.

Comments12 pages, to appear 42nd International Conference on Software Maintenance and Evolution

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑