发表机构
University of Calgary(卡尔加里大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出测试感知变异体生成方法,让LLM在提示中同时看到问题、解和基础测试,生成通过基础测试但被扩展测试捕获的真实缺陷,在HumanEval和MBPP上显著优于测试盲目提示和规则工具。
AI 中文摘要
变异测试通过向程序代码注入合成故障来评估测试套件的充分性。然而,传统的基于规则的工具通常会生成大量琐碎、冗余或等价的变异体,这限制了它们在识别测试套件缺口方面的实际应用。尽管最近基于大型语言模型(LLM)的方法能生成更真实的故障,但大多数方法仍是测试盲目的:模型仅看到源代码,无法推理现有测试已经覆盖的内容。我们提出了测试感知的变异体生成方法,其中LLM在单个提示中接收问题陈述、规范解和基础测试,并且必须生成一个通过基础单元测试的非平凡变异体。我们在HumanEval和MBPP基准上,对五组LLM——Gemini 3.1 Pro、Gemini 3 Flash、GPT 5.1 Codex Mini、GPT 4.1 Mini、Qwen3-32B——评估了这种方法。扩展的EvalPlus测试套件作为自动化预言机,用于验证存活的变异体是否代表真正的缺陷。测试感知提示生成的已验证故障率为87.7%(HumanEval)和79.1%(MBPP),这意味着这些变异体通过了所有基础测试但被预言机捕获。这大大优于匹配的测试盲目提示(分别仅产生12.2%和23.0%)以及传统的基于规则的工具mutmut(4.4%和5.7%)。虽然故障微妙性(变异体失败的扩展测试比例)在所有三种方法中保持相当,但测试感知相比测试盲目提示,最小化了每个已验证故障的计算成本。将现有单元测试暴露给LLM,使变异体生成从无目标的缺陷注入转向对现有测试套件弱点的有效发现。我们的工作为未来研究将测试感知变异体生成扩展到生产级环境奠定了坚实基础。
英文摘要
Mutation testing evaluates test-suite adequacy by injecting synthetic faults into program code. However, traditional rule-based tools often generate large numbers of trivial, redundant, or equivalent mutants that limit their practical use for identifying gaps in a test suite. While recent large language model (LLM)-based approaches generate more realistic faults, most remain test-blind: The model sees only the source code and cannot reason about what existing tests already cover. We propose test-aware mutant generation, in which an LLM receives the problem statement, canonical solution and base tests in a single prompt, and must generate a nontrivial mutant that passes the base unit tests. We evaluate this approach across a set of five LLMs -- Gemini 3.1 Pro, Gemini 3 Flash, GPT 5.1 Codex Mini, GPT 4.1 Mini, Qwen3-32B -- on the HumanEval and MBPP benchmarks. The extended EvalPlus test suites serve as an automated oracle to verify whether surviving mutants represent genuine bugs. Test-aware prompting yields verified fault rates of 87.7% (HumanEval) and 79.1% (MBPP), meaning these mutants pass all base tests but are caught by the oracle. This vastly outperforms the matched test-blind prompting (which yields only 12.2% and 23.0%, respectively) and the traditional rule-based tool mutmut (4.4% and 5.7%). While fault subtlety (the fraction of extended tests a mutant fails) remains comparable across all three methods, test-awareness minimizes the computational cost per verified fault, compared to test-blind prompting. Exposing an LLM to existing unit tests shifts mutant generation from untargeted bug injection toward effective discovery of weaknesses in an existing test suite. Our work establishes a concrete foundation for future research to scale test-aware mutant generation to production-level environments.
Comments20 pages, 7 figures. Accepted to ESEM 2026
DOI:10.4230/LIPIcs.ESEM.2026.49