arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

使用大语言模型(LLM)生成的回归测试的风险研究

On the Risks of using LLM-Generated Tests for Regression Testing

Mohammadali Charoosaei, Cedric Richter, Mike Papadakis

arXiv 2610.11835首次发表:更新:

发表机构

SnT, University of Luxembourg(卢森堡大学科学与技术研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对145个开源项目拉取请求,发现8%-17%的LLM生成回归测试为漏洞强化测试,未经人工验证会引发新型技术债务。

AI 中文摘要

软件处于持续迭代中:开发者不断添加功能、修复漏洞和重构代码,任何变更都可能破坏现有功能。回归测试通过在测试用例中捕获预期行为来防范此类影响。基于大语言模型(LLM)的测试生成旨在直接从被测代码自动生成回归测试,当实现正确时这很有益,但当代码存在漏洞时则会产生问题:生成的测试可能会编码并保留错误行为。为研究该风险,我们将基于LLM的回归测试生成应用于合并到软件项目主分支的拉取请求,并研究生成的测试对项目后续迭代的影响。我们区分两种测试:漏洞揭示测试(断言正确实现的行为)和漏洞强化测试(断言错误行为)。针对来自SciPy、Qiskit和pandas的145个拉取请求,8%-17%的生成测试为漏洞强化测试,仅2.4%-4.8%为漏洞揭示测试。漏洞强化测试会随时间持续存在:经过后续多次提交后,83%-91%的此类测试仍相关且可通过;它们还会累积:当所有拉取请求的漏洞合并到一个代码库中时,在提交历史末尾仍有83%-92%的漏洞被强化,而开发者编写的测试套件仅能检测到其中14%-30%的漏洞。我们的结果揭示了基于LLM的回归测试的一个根本风险:未经人工验证,它们可能将错误行为编码为预期行为,使漏洞在软件版本间持续存在并大幅规避开发者维护的测试套件,因此基于LLM的回归测试可能引发一种新型技术债务。

英文摘要

Software is under constant evolution: developers continuously add features, fix bugs, and refactor code, and any of these changes may break existing functionality. Regression testing guards against such effects by capturing expected behavior in test cases. LLM-based test generation aims to automate this process by generating regression tests directly from the code under test. This is beneficial when the implementation is correct, but problematic when the code contains faults: the generated tests may then encode and preserve incorrect behavior. To investigate this risk, we apply LLM-based regression test generation to pull requests merged into the main branch of software projects and study the impact of the generated tests on subsequent project evolution. We distinguish between fault-revealing tests, which assert correctly implemented behavior, and fault-enforcing tests, which assert faulty behavior. Across 145 pull requests from SciPy, Qiskit, and pandas, 8%-17% of the generated tests are fault-enforcing, while only 2.4%-4.8% reveal faults. Fault-enforcing tests persist over time: after several subsequent commits, 83%-91% of them are still relevant and pass. They also accumulate: when the faults of all pull requests are combined in one codebase, 83%-92% remain enforced at the end of the commit history, and the developer-written test suite detects only 14%-30% of them. Our results reveal a fundamental risk of LLM-generated regression tests: without manual validation, they may encode faulty behavior as expected behavior, allowing bugs to persist across software revisions and largely evade developer-maintained test suites. LLM-based regression testing can thus give rise to a new form of technical debt.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑