arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

故障感知与可解释的测试预言机预测

Fail-Aware and Explainable Test Oracle Prediction

Yue Zhao, Binish Tanveer, Jelena Zdravkovic

arXiv 2607.11342首次发表:更新:

发表机构

Department of Computer and Systems Sciences(计算机与系统科学系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对测试预言机构建难题,提出基于代码语言模型的FOCAL方法预测测试前缀成败,通过学习带标签数据、强调失败案例损失及基于行为证据预测,相比基线方法提升了未见过项目失败案例性能并提供更丰富解释。

AI 中文摘要

尽管测试预言机在故障检测中起着核心作用,但有效地构建它们仍然具有挑战性。最近基于学习的方法通过自动生成测试断言来应对这一挑战,但即便语法正确,这些断言通常也难以揭示错误。本研究探索了一种不同的方法,即训练一个模型直接预测给定的测试前缀是否会通过或失败。我们提出了FOCAL,一种基于新兴代码语言模型的判别式预言机预测器。它从带标签的测试前缀和被测方法对中学习,在训练期间采用强调失败情况的损失函数,并将其预测基于语句级行为证据。与基线方法SEER相比,我们在未见过的项目的失败案例上显著提高了性能,并提供了更丰富的解释。在故障检测基准和自动测试生成工件上的初步评估表明,我们的方法在其训练分布内具有高度准确性,并且在先前的判别式预言机失效的未见过的项目上显著提高了故障检测能力。此外,突出显示的语句得到行为解释检查的支持。这些早期结果表明,故障感知判别式预言机预测可以补充现有的方法,如模糊测试、基于搜索的测试和基于语言模型的测试生成。这些技术可以大规模生成测试前缀,但通常缺乏面向故障的预言机。在未来的工作中,FOCAL可以获取生成的测试前缀并为其附加故障感知预测的预言机,将大量输入生成转化为更有可能暴露语义故障的可执行测试。

英文摘要

Despite their central role in fault detection, test oracles remain challenging to construct effectively. Recent learning based methods address this challenge by automatically generating test assertions, yet even if syntactically correct, they are often ineffective in revealing bugs. Rather than generating assertions, this study explores a different approach by training a model to directly predict whether a given test prefix passes or fails. We present FOCAL, an emerging code LLM-based discriminative oracle predictor. It learns from labeled pairs of test prefixes and methods under test, employs losses that emphasize failing cases during training, and grounds its predictions in statement level behavioral evidence. Compared with the baseline method SEER, we substantially improve performance on failing cases for unseen projects and provide richer explanations. A preliminary evaluation on fault-detection benchmarks and automated test-generation artifacts shows that our approach is highly accurate within its training distribution and substantially improves failure detection on previously unseen projects where prior discriminative oracles collapse. Moreover, the highlighted statements are supported by behavioral explanation checks. These early results suggest that fail-aware discriminative oracle prediction can complement existing approaches such as fuzzing, search-based testing, and LLM-based test generation. These techniques produce test prefixes at scale but often lack fault oriented oracles. In future work, FOCAL could take generated test prefixes and attach fault-aware predicted oracles to them, turning high-volume input generation into executable tests that are more likely to expose semantic failures.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑