发表机构
Chengdu Institute of Computer Applications, Chinese Academy of Sciences; University of Chinese Academy of Sciences; Institute of Multidisciplinary Research for Advanced Materials (IMRAM), Tohoku University; Artificial Intelligence Research Institute, Shenzhen University of Advanced Technology(中国科学院成都计算机应用研究所; 中国科学院大学; 东北大学先进材料多学科研究所; 深圳先进技术研究院人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本综述围绕测试改变何种决策,系统梳理大语言模型与软件工程智能体中测试驱动方法,区分多种测试范式,提出机制分类法并指出研究方向。
AI 中文摘要
测试越来越多地参与大语言模型和软件工程智能体所做的决策。它们规定预期行为,指导程序构建与修复,选择候选方案,约束变换,并为软件分析提供执行证据。这些用途借鉴了测试驱动开发,但在测试顺序、预言可用性、可编辑工件以及执行的角色方面存在显著差异。我们提出了一项结构化的范围界定综述,围绕测试改变何种决策这一问题进行组织。该综述整合了87项研究及支持性记录,其中83项记录进行了方法级或协议级提取,并另外收集了五项实践资源。我们区分了红-绿-重构循环与测试条件生成、执行引导精化、测试中介分析和仅评估测试。然后,我们比较了代码生成、修复、翻译、重构、克隆检测、代码搜索、定位、训练数据构建和形式化规范验证。一项专门的分析考察了智能体工作流和可复用技能如何编码测试程序及其效果如何评估。在这些任务中,证据支持将测试可用性、测试有效性、反馈使用和评估独立性视为独立属性。仅通过测试并不能建立行为等价性、有效反馈或过程遵循;总体改进也可能掩盖不同模型、任务和分母之间的不同结果。我们将这些区别综合为一个机制分类法、一个跨任务比较和一个协议敏感的实证分析,并确定了预言验证、因果评估、长期维护和可复用测试驱动智能体能力方面的研究方向。
英文摘要
Tests increasingly participate in the decisions made by large language models and software engineering agents. They specify intended behavior, guide program construction and repair, select candidates, constrain transformations, and provide execution evidence for software analysis. These uses draw on test-driven development, yet differ substantially in test order, oracle availability, editable artifacts, and the role of execution. We present a structured scoping survey organized around the question of what decision a test changes. The review integrates 87 research and supporting records, with method- or protocol-level extraction for 83 records, alongside a separate collection of five practice resources. We distinguish the Red--Green--Refactor cycle from test-conditioned generation, execution-guided refinement, test-mediated analysis, and evaluation-only testing. We then compare code generation, repair, translation, refactoring, clone detection, code search, localization, training-data construction, and formal-specification validation. A dedicated analysis examines how agent workflows and reusable skills encode testing procedures and how their effects are evaluated. Across these tasks, the evidence supports treating test availability, test validity, feedback use, and evaluation independence as separate properties. Test passing alone does not establish behavioral equivalence, effective feedback, or process adherence; aggregate improvements can also conceal different outcomes across models, tasks, and denominators. We synthesize these distinctions into a mechanism taxonomy, a cross-task comparison, and a protocol-sensitive evidence analysis, and identify research directions in oracle validation, causal evaluation, long-horizon maintenance, and reusable test-driven agent capabilities