发表机构
York University(约克大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
NEUROTESTGEN结合符号执行与LLM,通过提取路径约束并引导测试生成,实现按需代码覆盖率,在多个LLM上显著优于现有方法。
AI 中文摘要
在自动化测试生成中,确保高结构覆盖率仍然是一个基本挑战,尤其是在复杂软件系统中,到达特定行或分支需要满足复杂的控制流和数据流约束。大语言模型(LLMs)最近在生成类人测试用例方面展现了强大的能力;然而,它们往往难以生成满足精确路径条件的输入。相反,符号执行可以系统地推导出这些约束,但它常常无法构造现实、可执行的测试用例,并受限于可扩展性问题。在本文中,我们介绍了NEUROTESTGEN,一种将符号执行与LLM驱动的测试合成相结合的混合方法,用于生成针对按需代码覆盖率的测试用例。给定方法内的一组目标语句,NEUROTESTGEN首先使用符号分析引擎(即Z3 SMT求解器)提取路径特定约束,并为期望的覆盖目标构建符号引导规范。然后,该规范用于引导LLM合成既结构有效又语义有意义的具体测试用例。对于涉及SMT求解器难以处理的复杂对象相关约束的路径,NEUROTESTGEN利用LLM推断合理的约束。此外,NEUROTESTGEN包含一个迭代反馈循环,用于验证LLM生成的测试并提供纠正性指导,直到覆盖目标行或分支或达到限制。我们在广泛使用的基准上的实证评估表明,NEUROTESTGEN在多种LLM(包括Llama 3.3 70B1、GPT-4o Mini、Claude 3.5 Haiku3和Claude Sonnet 4.6)上显著优于最先进的方法。
英文摘要
Ensuring high structural coverage remains a fundamental challenge in automated test generation, particularly for complex software systems where reaching specific lines or branches requires satisfying intricate control- and data-flow constraints. Large Language Models (LLMs) have recently demonstrated strong capabilities in producing human-like test cases; however, they often struggle to generate inputs that satisfy precise path conditions. Conversely, symbolic execution can systematically derive such constraints, but it often fails to construct realistic, executable test cases and is constrained by scalability limitations. In this paper, we introduce NEUROTESTGEN, a hybrid approach that integrates symbolic execution with LLM-driven test synthesis to generate test cases targeting on-demand code coverage. Given a set of target statements within a method, NEUROTESTGEN first employs a symbolic analysis engine (i.e., the Z3 SMT solver) to extract path-specific constraints and construct a symbolic guidance specification for the desired coverage goal. This specification is then used to guide an LLM in synthesizing concrete test cases that are both structurally valid and semantically meaningful. For paths involving complex object-related constraints that are difficult for SMT solvers to handle, NEUROTESTGEN leverages LLMs to infer plausible constraints. Furthermore, NEUROTESTGEN incorporates an iterative feedback loop that validates LLM-generated tests and provides corrective guidance until the target line or branch is covered or a limit is reached. Our empirical evaluation on a widely used benchmark demonstrates that NEUROTESTGEN significantly outperforms the state-of-the-art approach across multiple LLMs, including Llama 3.3 70B1, GPT-4o Mini, Claude 3.5 Haiku3, and Claude Sonnet 4.6.