发表机构
Delft University of Technology(代尔夫特理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出LLMSuite混合测试框架,结合自我改进提示与类级LLM推理,增强NLP库的自动化单元测试生成,显著提升覆盖率与突变得分。
AI 中文摘要
自动化单元测试生成工具(如EvoSuite)在通用软件上表现良好,但在处理领域特定软件(如自然语言处理(NLP)库)时往往面临困难,因为这类软件的输入必须遵循语义、句法和结构约束。大型语言模型(LLMs)能够生成与领域相关的测试代码,但仅由LLMs生成的测试往往无法编译或达到足够的覆盖率。我们提出了LLMSuite,一种混合测试生成框架,它将自我改进提示与类级LLM推理集成到基于搜索的测试过程中。在该机制中,LLM根据先前生成的反馈迭代改进其测试片段。这使得模型能够生成越来越精确、与领域一致的代码片段,从而引导进化搜索朝向执行复杂且难以触及的行为。当底层进化算法中多个代际没有目标改进时,这些改进后的片段会被解析并注入EvoSuite的种群中以扩展搜索空间。为支持我们的评估,我们构建了一个新数据集,包含来自五个广泛使用的Java NLP项目的100个类。我们还用Java重新实现了CodaMOSA(一种最新的混合SBST-LLM技术),以便进行直接比较。在该数据集上,LLMSuite的分支覆盖率和行覆盖率分别提高了约10%和8%,并且突变得分比CodaMOSA-J高11%。与EvoSuite相比,LLMSuite的分支覆盖率和行覆盖率提高了约15%,突变得分提高了5%。与仅使用LLM的基线相比,它将结构覆盖率提高了36%,突变得分提高了约24.7个百分点。最后,LLMSuite通过执行通常未被测试的领域特定行为,补充了手动编写的测试套件。
英文摘要
Automated unit test generation tools like EvoSuite perform well on general-purpose software but often struggle with domain-specific software such as Natural Language Processing (NLP) libraries, where inputs must follow semantic, syntactic, and structural constraints. Large Language Models (LLMs) can generate domain-relevant test code, but tests produced by LLMs alone often fail to compile or achieve sufficient coverage. We propose LLMSuite, a hybrid test generation framework that integrates self-refinement prompting with class-level LLM reasoning into the search-based testing process. In this mechanism, the LLM iteratively improves its test snippets based on feedback from previous generations. This enables the model to produce increasingly precise, domain-consistent code fragments that steer the evolutionary search toward exercising complex and otherwise hard-to-reach behaviors. When no objective improves over multiple generations in the underlying evolutionary algorithm, these refined snippets are parsed and injected into EvoSuite's population to expand the search space. To support our evaluation, we constructed a new dataset comprising 100 classes drawn from five widely used Java NLP projects. We also re-implemented CodaMOSA, a recent hybrid SBST-LLM technique, in Java to enable a direct comparison. Across this dataset, LLMSuite improves branch and line coverage by approximately 10% and 8%, respectively, and achieves an 11% higher mutation score than CodaMOSA-J. Compared to EvoSuite, LLMSuite yields roughly 15% higher branch and line coverage and 5% higher mutation score. Against an LLM-only baseline, it improves structural coverage by 36% and mutation score by about 24.7 percentage points. Finally, LLMSuite complements manually written test suites by exercising domain-specific behaviors that are often left untested.
Comments25 pages, 8 figures, 8 tables