发表机构
University of Toronto(多伦多大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型生成测试套件时,覆盖度和变异分数与有效性的关系。通过重复研究发现其有用性依赖上下文,在不同场景表现不同,且测试套件大小非主要混杂因素,为评估此类测试生成提供指导。
AI 中文摘要
大语言模型(LLMs)的进展引发了使用其自动化测试生成的兴趣。先前工作常用代码覆盖度和变异分数等代理指标评估生成的测试套件。然而,已有研究表明人工编写测试中,控制测试套件大小后,覆盖度、变异和实际错误检测间的相关性大多消失。本文对两项先前工作进行大规模重复研究,重新审视覆盖度、变异和实际错误检测有效性间的关系。结果表明覆盖度和变异的有用性高度依赖上下文:在回归风格设置中,这些指标在跨模型比较时能提供有意义信号;在被测代码可能有错误的场景下,它们不再是可靠指标。且几乎没有证据表明测试套件大小是大语言模型生成测试中覆盖度、变异和实际错误检测相关性的主要混杂因素。基于这些发现,讨论了如何解释先前研究结果并为评估基于大语言模型的测试生成提供可行指导。
英文摘要
Recent advances in large language models (LLMs) have driven growing interest in using LLMs to automate test generation. Prior work commonly evaluates generated test suites using proxy metrics such as code coverage and mutation score. However, studies by Inozemtseva et al. and Papadakis et al. show that, for human-written tests, correlations among coverage, mutation, and real-bug detection can largely vanish once test suite size is controlled, raising concerns about the validity of evaluations based on proxy metrics. It also remains unclear whether these conclusions carry over to LLM-generated tests, given that prevailing LLM-based test-generation workflows differ substantially from traditional approaches. In this paper, we conduct a large-scale replication study of these two prior works using a wide range of test suites generated by a diverse set of LLMs, and re-examine the relationships among coverage, mutation, and real-bug detection effectiveness. Our findings diverge substantially from prior results. We show that the usefulness of coverage and mutation is highly context-dependent: in regression-style settings where the code provided to the LLM can be reasonably assumed bug-free, these metrics can provide meaningful signals when comparing across models; in another common scenario where the code-under-test may already be buggy and the goal is to expose the bug within the code-under-test, they no longer serve as reliable indicators. We also find little evidence that test suite size is a dominant confounder for correlations among coverage, mutation, and real-bug detection for LLM-generated tests. Based on these findings, we discuss how to interpret results from prior studies and provide actionable guidance for evaluating LLM-based test generation.
Comments24 pages, 4 figures, 14 tables. Accepted at ISSTA 2026; to appear in Proceedings of the ACM on Software Engineering (PACMSE), Vol. 3, No. ISSTA, Article ISSTA002