arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

评估类级LLM生成的Python测试套件的有效性

Evaluating the effectiveness of class-level LLM-generated test suites in Python

Bilal Al-Ahmad, M. Harshvardhan, Khaled El-Fakih, Anas AlSobeh

arXiv 2609.24341首次发表:更新:

发表机构

American University of Sharjah; The University of Jordan; Utah Valley University(沙迦美国大学; 约旦大学; 犹他谷大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究评估提示策略和模型选择对LLM生成Python测试套件的影响,发现可执行性差异大,模型选择比提示更重要,建议以可执行性为门槛并结合变异测试评估。

AI 中文摘要

背景:大型语言模型(LLM)能够快速生成单元测试,但高结构覆盖率并不能证明这些测试能够可靠执行或检测故障。现有证据通常将覆盖率作为主要结果,很少通过类级变异测试来比较提示策略和模型。目标:本研究考察提示策略和模型选择如何影响LLM生成的Python测试套件相对于人工编写的测试套件的可执行性、结构覆盖率、故障检测有效性和结构质量。方法:我们在ClassEval基准上,针对多种当前LLM配置,评估了多种提示策略。评估结合了执行结果、行和分支覆盖率、Cosmic Ray变异得分以及结构质量指标。主要分析将成功执行视为先决条件;配对比较仅使用相关可执行子集共有的类。结果:结构覆盖率始终接近其上限,在配置之间几乎没有区分度。可执行性差异显著。所提出的提示在变异得分上表现强劲,但没有一个提示在所有模型上占优。模型选择比提示选择解释了更多的变异,并且它们的交互表明提示的有效性取决于所选模型。人工和LLM测试套件在不相等的可执行子集上进行评估,因此它们的相对变异得分不能确立优越性。结论:对LLM生成测试的可靠评估应将可执行性作为门槛,并将覆盖率与变异测试和结构质量指标相结合。在实践中,模型选择应先于提示调优。

英文摘要

Context: Large language models (LLMs) can generate unit tests quickly, but high structural coverage does not establish that those tests execute reliably or detect faults. Existing evidence often treats coverage as the principal outcome and rarely compares prompt strategies and models through mutation testing at class level. Objective: This study examines how prompt strategy and model choice shape the executability, structural coverage, fault-detection effectiveness, and structural quality of LLM-generated Python test suites relative to human-written suites. Method: We evaluate multiple prompt strategies across a diverse set of current LLM configurations on the ClassEval benchmark. The evaluation combines execution outcomes, line and branch coverage, Cosmic Ray mutation scores, and structural quality indicators. Primary analyses treat successful execution as a prerequisite; paired comparisons use only classes shared by the relevant executable subsets. Results: Structural coverage is consistently near its ceiling and offers little discrimination among configurations. Executability varies substantially. The proposed prompt performs strongly for mutation score, but no prompt dominates across models. Model choice explains more variation than prompt choice, and their interaction shows that prompt effectiveness depends on the selected model. Human and LLM suites are evaluated on unequal executable subsets, so their relative mutation scores do not establish superiority. Conclusion: Reliable assessment of LLM-generated tests should treat executability as a gate and combine coverage with mutation testing and structural quality indicators. In practice, model selection should precede prompt tuning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑