arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

评估和减轻大语言模型生成单元测试中错误代码的误导效应

Evaluating and Mitigating the Misguidance Effect of Buggy Code in LLM-Generated Unit Tests

Junda Zhao, Shurui Zhou, Eldan Cohen

arXiv 2607.22883首次发表:更新:

发表机构

University of Toronto(多伦多大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究大语言模型生成单元测试时错误代码的误导效应,提出新指标衡量,分析其双重影响,引入基于规范的单元测试生成范式,有效减少误导性测试,增加有效测试,改善测试生成管道,适用于各类代码。

AI 中文摘要

虽然大语言模型在自动化单元测试生成方面前景广阔,但近期研究表明,用错误代码提示模型时,生成测试的质量会受到负面影响。本文提出一种新指标来定量衡量‘误导效应’,即错误代码引导大语言模型生成验证其错误行为而非暴露错误的测试。分析显示,用错误代码提示大语言模型有严重的双重影响:显著增加断言错误行为的‘误导性测试’,同时抑制有效找错测试的生成。从模型内部角度进一步证实了这种效应。为应对此,引入并验证了基于规范的单元测试生成范式,用大语言模型生成的规范文档字符串替换提示中的被测代码。结果表明,该范式有效减少误导性测试,大幅增加有效测试,改善多轮反馈驱动的测试生成管道,且适用于有错误和无错误的代码。总体而言,基于规范的提示是减轻大语言模型生成单元测试中错误代码误导的有前景策略。

英文摘要

While Large Language Models (LLMs) show great promise for automating unit test generation, recent studies suggest that the quality of generated tests can be negatively impacted when models are prompted with buggy code. This paper presents a new metric to quantitatively measure the "misguidance effect," a phenomenon where buggy code steers LLMs toward generating tests that validate its erroneous behavior rather than expose it. Our analysis reveals that prompting LLMs with buggy code has a severe, twofold impact: it significantly increases "misguided tests" that assert incorrect behavior while simultaneously suppressing the generation of effective, bug-finding tests. We further corroborate this effect from a model-internal perspective, showing that buggy code skews LLMs' preference toward tests that assert the same erroneous behavior. To counter this, we introduce and validate a specification-based unit test generation paradigm that replaces the code under test in the prompt with an LLM-generated specification docstring. Our results show that this paradigm effectively reduces misguided tests while substantially increasing effective tests, improves multi-round, feedback-driven test generation pipelines, and remains applicable to both buggy and bug-free code. Overall, these results suggest that specification-based prompting is a promising strategy for mitigating misguidance from buggy code in LLM-generated unit tests.

Comments24 pages, 7 figures, 12 tables. Accepted at ISSTA 2026; to appear in Proceedings of the ACM on Software Engineering (PACMSE), Vol. 3, No. ISSTA, Article ISSTA113

DOI:10.1145/3832204 10.1145/3832204

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑