arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15188cs.SEcs.AI

Claude AI生成的Python测试质量并不逊于人类编写的测试

The Quality of Claude AI-authored Python Tests Is Not Weaker Than Human-authored Tests

Douglas J. Leith

首次发表
浏览论文内容

中文总结 AI 辅助

该研究对比Django、Pandas的人类与Claude AI生成的Python测试,采用多协议评分,发现最新Claude模型的测试质量不逊于人类,且研究设置更贴合真实场景。

中文摘要 AI 辅助

我们针对两个成熟开源项目Django和Pandas,将Claude AI编写的Python测试与人类编写的Python测试的质量进行了评估。每个语料库各有数百个测试,均在同一协议下评分。使用单侧非劣效界值,我们发现最新的Claude模型(Sonnet/Opus 4.6及更高版本)编写的测试并不比这两个人类编写的语料库差。本研究具有三点特征:其一,AI编写的语料库来自真实工具的测试,而非其他AI测试生成研究中采用的针对固定目标孤立生成的合成测试;其二,每个测试均依据三个独立的故障注入协议以及七维度定性设计标准单独评分,使各方法可交叉验证;其三,测试以单个而非套件级评分,能精准识别需关注的具体测试。

英文摘要

We evaluate the quality of Claude AI-written Python tests against human-written Python tests from two established open-source projects Django and Pandas. Hundreds of tests per corpus are scored under one identical protocol. Using one-sided non-inferiority bounds, we find that the tests written by recent Claude models (Sonnet/Opus 4.6 and later) are no weaker than the two human-written corpora. In this study: (i) the AI-written corpus is tests from real tools, not synthetic tests generated in isolation against a fixed target, the setup used by every other AI-test-generation study we are aware of; (ii) every test is individually scored under three independent fault-injection protocols plus a seven-axis qualitative design rubric, allowing methods to cross-validate each other; (iii) tests are scored individually, rather than suite-level, identifying exactly which specific tests need attention.

发表机构

  • Trinity College Dublin(都柏林圣三一学院)

机构由 AI 辅助整理,请以论文原文为准。

↑