arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于大语言模型的测试生成:信息源、生成策略与质量证据

LLM-Based Test Generation: Information Sources, Generation Strategies, and Quality Evidence

Yunhao Liang, Chengguang Gan, Ruixuan Ying, Hanjun Wei, Zhe Cui, Shiwen Ni

arXiv 2610.05001首次发表:更新:

发表机构

Chengdu Institute of Computer Applications, Chinese Academy of Sciences; University of Chinese Academy of Sciences; Institute of Multidisciplinary Research for Advanced Materials (IMRAM), Tohoku University; Artificial Intelligence Research Institute, Shenzhen University of Advanced Technology(中国科学院成都计算机应用研究所; 中国科学院大学; 东北大学先进材料多学科研究所; 深圳先进技术研究院人工智能研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文通过结构化综述,基于95条来源记录,从四个维度梳理大语言模型生成测试的方法,阐明不同测试证据类型,并提出独立预言评估、预算感知生成等研究议程。

AI 中文摘要

大语言模型越来越多地被用于生成测试场景、可执行测试套件、断言、交互序列以及模糊测试基础设施。这些工件服务于不同的目的,并依赖于关于正确行为的不同信息源。因此,增加实现覆盖率的测试、与参考程序一致的测试以及检测需求违规的测试提供了不同类型的证据。我们提出了一项结构化的叙述性综述,将生成过程与用于证明测试质量的证据联系起来。基于截至2026年9月29日精选的95条来源记录,我们沿四个维度组织该领域:测试目标与工件、训练和生成期间可用的信息、生成与学习机制以及评估证据。我们综合了单元测试和基于需求的测试、测试预言、API和GUI测试、模糊测试以及学习型测试生成器方面的工作。由此产生的框架解释了执行反馈如何提高可执行性,同时影响预言的独立性;为何规范来源和训练时参考监督对比较很重要;以及下游代码选择增益如何不同于测试正确性。我们利用这些区别来组织基准,并推导出一项研究议程,涵盖独立预言评估、预算感知生成、仓库规模评估和测试维护。

英文摘要

Large language models are increasingly used to generate test scenarios, executable test suites, assertions, interaction sequences, and fuzzing infrastructure. These artifacts serve different purposes and rely on different sources of information about correct behavior. A test that increases implementation coverage, a test that agrees with a reference program, and a test that detects a requirements violation therefore provide distinct kinds of evidence. We present a structured narrative survey that connects the generation process to the evidence used to justify test quality. Drawing on 95 curated source records through 29 September 2026, we organize the field along four dimensions: testing objectives and artifacts, information available during training and generation, generation and learning mechanisms, and evaluation evidence. We synthesize work on unit and requirements-based testing, test oracles, API and GUI testing, fuzzing, and learned test generators. The resulting framework explains how execution feedback can improve executability while also affecting the independence of an oracle, why specification provenance and training-time reference supervision matter for comparisons, and how downstream code-selection gains differ from test correctness. We use these distinctions to organize benchmarks and derive a research agenda for independent oracle assessment, budget-aware generation, repository-scale evaluation, and test maintenance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑