超越测试存在:评估开源项目中代理生成测试的质量和稳健性
Beyond Test Presence: Assessing the Quality and Robustness of Agent-Generated Tests in Open-Source Projects
浏览论文内容
中文总结 AI 辅助
研究开源项目中代理生成测试的质量和稳健性,通过对大量测试工件进行实证比较,用“白盒”静态分析框架评估质量维度,发现代理在边缘情况覆盖率上优于人类,但生成测试更易脆弱,揭示其缺乏编写稳定测试的“环境意识”。
中文摘要 AI 辅助
人工智能驱动的编码代理集成到持续集成/持续交付(CI/CD)管道中,从根本上改变了软件验证的方式。当前评估基准主要关注通过率而非生成测试的内在质量,可能导致“隐性技术债务”。我们通过对204,673个测试工件进行大规模实证比较来解决这一方法差距,这些工件包括来自AIDev数据集的24,941个人工编写文件和179,732个代理生成文件。使用Python的朴素ast模块进行抽象语法树(AST)解析,实现了一个“白盒”静态分析框架来评估三个质量维度:断言强度(RQ1)、边缘情况覆盖率(RQ2)和脆弱性潜力(RQ3)。结果显示,人工智能代理在边缘情况覆盖率方面优于人类,但生成的测试存在更高的脆弱性风险。这些发现表明,人工智能代理在严格的边界测试方面表现出色,但缺乏编写稳定、封闭测试所需的“环境意识”。
英文摘要
The integration of AI-powered coding agents into Continuous Integration/Continuous Delivery (CI/CD) pipelines has fundamentally altered how software verification is conducted. While these agents successfully automate the test generation, current evaluation benchmarks (e.g., SWE-bench) largely focus on pass-rates rather than the intrinsic quality of the generated tests. This raises the possibility of "stealth technical debt", in which test suites pass execution but do not offer comprehensive coverage or semantic value. We address this methodological gap through a large-scale, empirical comparison of 204,673 test artifacts which comprises of 24,941 human-authored files and 179,732 agent-generated files; sourced from the AIDev dataset. Using the Abstract Syntax Tree (AST) parsing with Python's naive ast module, we implemented a "white-box" static analysis framework to evaluate three quality dimensions: Assertion Strength (RQ1), Edge-Case Coverage (RQ2), and Flakiness Potential (RQ3). Our results present a nuanced inversion of traditional assumptions. AI agents performed better than humans in Edge-Case Coverage, with almost twice the variety of boundary checks (Variety Score: 0.62 vs 0.32) and a higher frequency of null-safety testing (13.40% vs. 8.3%), even though human developers had a slight advantage in Assertion Strength (88.1% strong assertions vs. 85.37% for agents). But this thoroughness comes at a price: due mostly to their reliance on file I/O and non-deterministic logic, agent-generated tests exhibited a higher risk of flakiness (Candidate Rate: 0.41 vs. 0.30). These findings suggest that while AI agents excel at rigorous boundary testing, they lack the "environmental awareness" needed to write stable, hermetic tests.