arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.18711cs.SE

基于大语言模型的软件功能缺陷不变式测试

LLM-Based Invariant Testing for Software Functional Bugs

Ruogu Yang, Yifeng He, Yundi Xu, Yuqing Wei, Hao Chen

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对软件功能缺陷,提出基于大语言模型的LISA不变式测试框架,在API n-gram反馈指导下迭代生成API序列和程序不变式,相比其他方法有更高缺陷检测率和代码覆盖率,能为开发者提供高可信度缺陷候选。

中文摘要 AI 辅助

手动编写单元测试来发现软件库中的功能缺陷既耗时又需要深入理解API的预期语义。基于启发式的测试生成方法可用性低,因为它们无法像人类一样推理程序语义或解释源代码和文档。传统的模糊测试技术如OSS-Fuzz通常依赖崩溃来检测缺陷,但功能缺陷并不总是导致崩溃。为克服这些限制,我们提出了LISA,一种用于软件功能缺陷的新型基于大语言模型的不变式测试框架。LISA在API n-gram反馈的指导下迭代生成API序列和程序不变式,与模糊测试和先前基于大语言模型的测试生成方法相比,具有更高的缺陷检测率和有竞争力的代码覆盖率,并将每个发现报告为高可信度的缺陷候选供开发者确认。

英文摘要

Manually writing unit tests to uncover functional bugs in software libraries is not only time-consuming but also requires a deep understanding of the intended semantics of the APIs. Heuristic-based test generation methods suffer from low usability because they cannot reason about program semantics or interpret source code and documentation as humans do. Traditional fuzzing techniques like OSS-Fuzz often rely on crashes to detect bugs, but functional bugs do not always cause crashes. To overcome these limitations, we present LISA, a novel LLM-based invariant testing framework for software functional bugs. LISA iteratively generates API sequences and program invariants guided by API n-gram feedback, achieving higher bug-detection rates and competitive code coverage compared with both fuzzing and prior LLM-based test generation approaches, and reporting each finding as a high-confidence bug candidate for developer confirmation.

补充信息

↑