arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从业务需求到测试断言:评估大语言模型生成的预言机在实际错误上的表现

From Business Requirements to Test Assertions: Evaluating LLM-Generated Oracles on Real Bugs

Tiancheng Ma, Nasir U. Eisty

arXiv 2607.10277首次发表:更新:

AI 中文总结

研究大语言模型能否从自然语言业务需求生成测试预言机,提出基于Defects4J的流程,对10个实际错误进行实验,评估预言机在与需求衍生预言机及被测系统一致性上表现,发现大语言模型有泛化但存在差异,为后续研究提供可行性参考。

AI 中文摘要

预言机问题(确定测试的正确预期结果)仍是自动化测试的主要瓶颈,随着非专家依赖无法可靠验证的人工智能生成代码,该问题愈发重要。我们研究大语言模型能否在不访问源代码或示例输入输出对的情况下,直接从自然语言业务需求生成可推广的测试预言机。我们提出了一个基于Defects4J的可重现、需求驱动的流程。对于Defects4J Lang中的10个实际错误,我们提取行为变化,手动将变化转化为业务需求,构建需求衍生预言机作为黄金标准,并促使五个大语言模型生成Java预言机代码。我们在两个目标下评估预言机的正确性和泛化能力,报告宏观平均准确率、精确率、召回率和F1值。大语言模型实现了一定程度的泛化,但在错误和模型层面存在显著差异。生成的预言机与需求衍生预言机的一致性比与被测系统的一致性更高,需求技术性/模糊性评级与预言机准确性之间的相关性较弱且置信区间较宽。在该数据集中,需求属性与预言机准确性之间不存在可检测的线性关系,这表明预训练覆盖范围和所需行为的语义特异性主导预言机正确性。作为概念验证的初步证明,这些发现是初步的,旨在确立可行性并推动更大规模的实证研究。

英文摘要

The oracle problem (determining the correct expected outcome for a test) remains a major bottleneck in automated testing, and is increasingly relevant as non-experts rely on AI-generated code they cannot reliably validate. We study whether large language models (LLMs) can generate generalizable test oracles directly from natural-language business requirements, without access to source code or example input-output pairs. We propose a reproducible, requirement-driven pipeline grounded in Defects4J. For each of 10 real bugs from Defects4J Lang (Bugs 1 and 3-11), we (i) extract behavioral changes via buggy/fixed diffs, (ii) manually translate the change into a business requirement, (iii) construct a requirement-derived oracle (REQ) as a gold standard, and (iv) prompt five LLMs (DeepSeek-V3, Gemma-3n, Llama-3, Mistral-7B, and Qwen-3) to generate Java oracle code. We evaluate oracle correctness and generalization under two targets: agreement with REQ and agreement with the system under test (SUT), reporting macro-averaged accuracy, precision, recall, and F1. LLMs achieve non-trivial generalization but with substantial bug- and model-level variance. Generated oracles align more closely with REQ than with SUT, and correlations between requirement technicality/ambiguity ratings and oracle accuracy are weak with wide confidence intervals. No detectable linear relationship exists between requirement properties and oracle accuracy in this dataset, suggesting that pretraining coverage and the semantic specificity of the required behavior dominate oracle correctness. As a pilot proof of concept, these findings are preliminary and are intended to establish feasibility and motivate larger-scale empirical investigation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑