arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

传统测试准则在检测大语言模型生成代码中的缺陷方面效果如何?

How effective are traditional test criteria at detecting bugs in large language models generated code?

Asma Hamidi, Michael Konstantinou, Renzo Degiovanni, Mike Papadakis

arXiv 2609.09315首次发表:更新:

发表机构

SnT, University of Luxembourg; Luxembourg Institute of Science and Technology(卢森堡大学 SnT; 卢森堡科学与技术研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过实证评估5个LLM和4个基准中的6000多个故障实例,发现传统测试准则对LLM生成代码的缺陷检测率极低,变异测试优势有限,需手动推理断言。

AI 中文摘要

测试充分性准则被广泛用于评估和指导软件测试。尽管先前的研究已使用人工编写的程序、故障和测试对这些准则进行了广泛考察,但大语言模型(LLMs)在代码生成中的日益普及引发了关于这些准则在检测LLM引发的故障方面有效性的重要问题。为探究此问题,我们开展了一项实证研究,涉及5个LLM和4个基准,模拟了代码和测试均自动生成的端到端工作流。我们收集了6000多个故障程序实例,并评估了3种广泛使用的充分性准则的有效性和效率:语句覆盖、分支覆盖和变异测试。我们的研究结果揭示了几个关键见解。首先,LLM引入的大多数故障相对容易捕获。其次,具有挑战性的故障难以通过传统的基于覆盖或基于变异的准则触发。第三,实际故障检测率仍然极低,通常接近零,因为测试预言未能捕获由生成的测试前缀触发的故障行为,这暴露了自动化测试生成的关键局限性。第四,提示感知的预言可以改善故障检测,但其整体有效性仍然有限,凸显了用户需要手动推理测试断言的必要性。我们进一步观察到,变异测试在触发和检测故障方面仅略优于传统覆盖准则,这引发了对其在此情境下显著更高的应用成本是否合理的问题。

英文摘要

Test adequacy criteria are widely used to evaluate and guide software testing. Although prior research has extensively examined these criteria using human-written programs, faults, and tests, the increasing adoption of Large Language Models (LLMs) for code generation raises important questions about their effectiveness in detecting LLM-induced faults. To investigate this, we conduct an empirical study involving 5 LLMs and 4 benchmarks, simulating end-to-end workflows in which both code and tests are automatically generated. We collect 6,000+ faulty program instances and evaluate the effectiveness and efficiency of 3 widely used adequacy criteria: statement coverage, branch coverage, and mutation testing. Our findings reveal several key insights. First, most faults introduced by LLMs are relatively trivial to catch. Second, the challenging faults are difficult to trigger using either traditional coverage-based or mutation-based criteria. Third, actual fault detection rates remain extremely low, often near zero, because test oracles fail to capture faulty behavior triggered by the generated test prefixes, exposing a critical limitation of automated test generation. Fourth, prompt-aware oracles can improve fault detection, but their overall effectiveness remains limited, highlighting the need for users to manually reason about test assertions. We further observe that mutation testing only marginally outperforms traditional coverage criteria in both triggering and detecting faults, raising questions about whether its significantly higher application cost is justified in this context.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑