文档与代码模式:什么驱动基于大语言模型的异常预言生成?
Documentation vs. Code Patterns: What Drives LLM-Based Exception Oracle Generation?
浏览论文内容
中文总结 AI 辅助
该研究通过干预式与归因引导的替换消融实验,发现基于大语言模型的异常预言生成(TOG)的高准确率多源于捷径信号而非结构化异常文档,挑战了准确率反映语义稳健性的假设,提出需从证据依据角度评估TOG系统。
中文摘要 AI 辅助
基于大语言模型的测试预言生成(TOG)方法在异常预言生成任务上报告了较高的准确率,但目前仍不清楚是什么证据驱动了这些预测。具体而言,模型是使用Javadoc @throws子句这类显式的异常行为文档,还是依赖测试、代码和文档中的重复模式?我们通过一项大规模的干预式研究来调查这个问题,研究涵盖了三种TOG系统,包括基于分类器和生成式架构,模型规模从约1.1亿到70亿参数,在三个真实基准上进行评估,其中包括两个生成测试数据集和一个新的开发者编写测试基准。我们首先移除Javadoc @throws子句,发现准确率仅发生微小变化,最大降幅低于1个百分点,这表明结构化异常文档不是异常预言预测的主要驱动因素。随后,我们应用归因引导的替换消融来识别预测所依赖的信号。结果显示,高准确率可能由捷径信号驱动:一些模型对少量结构标记高度敏感,而另一些模型则将依赖分布在多个词汇线索上。这些发现挑战了“高异常预言准确率反映对异常语义的稳健使用”这一假设。因此,未来的TOG系统不仅应通过是否预测出正确的预言类型来评估,还应通过其预测是否基于有意义的异常触发证据来评估。
英文摘要
LLM-based test oracle generation (TOG) methods report high accuracy on exception oracle generation, but it remains unclear what evidence drives these predictions. In particular, do models use explicit exceptional-behavior documentation such as Javadoc @throws clauses, or do they rely on recurring patterns in tests, code, and documentation? We investigate this question through a large-scale intervention-based study of three TOG systems spanning classifier-based and generative architectures and model sizes from roughly 110M to 7B parameters, evaluated on three real-world benchmarks comprising two generated-test datasets and a new benchmark of developer-written tests. We first remove Javadoc @throws clauses and find that accuracy changes only marginally, with the largest drop below one percentage point. This indicates that structured exception documentation is not the primary driver of exception-oracle prediction. We then apply attribution-guided substitution ablations to identify the signals that predictions depend on. The results show that high accuracy can be driven by shortcut signals: some models are highly sensitive to a small number of structural tokens, while others distribute reliance across many lexical cues. These findings challenge the assumption that strong exception-oracle accuracy reflects robust use of exception semantics. Future TOG systems should therefore be evaluated not only by whether they predict the correct oracle type, but also by whether their predictions are grounded in meaningful exception-triggering evidence.