AI 中文总结
研究如何将历史真实错误机制转化为可执行反馈目标用于基于大语言模型的单元测试生成,提出相应框架,在Defects4J任务上评估,结果显示该机制引导的合成错误反馈能提高真实错误检测率,是引导测试的有效方式。
AI 中文摘要
大语言模型为单元测试生成带来新机遇,但可执行测试不一定能揭示实际缺陷。本文研究如何将历史真实错误机制转化为基于大语言模型的单元测试生成的可执行反馈目标。所提框架构建真实错误记录的结构和语义表示,检索适用于焦点方法的机制并实例化为合成错误以指导迭代测试增强。在Defects4J的方法级真实错误检测任务上评估该方法,结果表明机制引导的合成错误反馈比基于执行、覆盖、变异、知识和搜索的基线更能提高真实错误检测率。结果表明,将真实错误机制组织为可检索和可执行的反馈目标是引导生成测试朝向错误触发输入和行为预言的有效方法。
英文摘要
Large language models (LLMs) have opened new opportunities for unit test generation, but executable tests do not necessarily reveal real defects. This paper studies how historical real-bug mechanisms can be transformed into executable feedback targets for LLM-based unit test generation. The proposed framework constructs structural and semantic representations of real-bug records, retrieves mechanisms applicable to a focal method, and instantiates them as synthetic bugs that guide iterative test enhancement. We evaluate the approach on method-level real-bug detection tasks from Defects4J and show that mechanism-guided synthetic-bug feedback improves real-bug detection over execution-, coverage-, mutation-, knowledge-, and search-based baselines. The results suggest that organizing real-bug mechanisms as retrievable and executable feedback targets is an effective way to guide generated tests toward bug-triggering inputs and behavioral oracles.
Comments12 pages, 7 figures, 6 tables. Preprint