发表机构
Singapore Management University; Monash University; GovTech(新加坡管理大学; 蒙纳士大学; 政府科技局)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
VulContextBench通过评估编码智能体是否检索到其结论依赖的代码而非结论本身,揭示了前沿模型在探索与报告证据间的巨大差距,填补了现有漏洞检测基准的盲区。
AI 中文摘要
漏洞检测基准评估的是智能体得出的结论,而非其收集的证据。因此,一个从预训练中回忆起某个CVE的模型,与一个追踪了数据流的模型得分相同。我们研究了一个这种差异至关重要的任务:判断一次提交是否引入了漏洞。我们不对结论进行评分,而是评估智能体是否检索到了其结论所依赖的代码。我们提出了VulContextBench,一个包含111个漏洞引入提交(VICs)、跨越83个代码仓库、63个CWE和5种语言的基准。现有数据集通过将修复追溯到版本历史来标记此类提交,这往往指向错误的提交。因此,我们根据VIC的明确四标准定义,对每个案例进行人工审计,使基准不继承这种标签噪声。每个案例都标注了黄金上下文,共464个代码块,每个代码块都按其在该漏洞证据中的角色进行了标记。我们使用精确率、召回率和F1分数在三个粒度(文件、代码块和行)上评估了七个前沿模型,并分别对智能体在探索过程中查看的上下文和其最终声明为证据的上下文进行评分。两者之间的差距是主要发现。每个模型在探索时都打开了大部分黄金上下文,但只将其中一部分报告为证据。在代码块级别,模型报告的黄金上下文占比比其查看的占比低37到73个百分点。Qwen3-Coder-Next查看了标注代码块中86.3%的行,但在最终报告中仅引用了12.9%。引用最多的GPT-5.5查看了73%并报告了36%。这些结果凸显了找到相关代码与将其选择到最终报告之间的差距,而结论级别的基准无法揭示这一点。
英文摘要
Vulnerability-detection benchmarks score the verdict an agent reaches, not the evidence it gathered. A model that recalls a CVE from pretraining therefore scores the same as one that traced the data flow. We study a task where this difference matters, deciding whether a commit introduces a vulnerability. Instead of scoring the verdict, we score whether the agent retrieved the code its conclusion depends on. We present VulContextBench, a benchmark of 111 vulnerability-introducing commits (VICs) across 83 repositories, 63 CWEs, and five languages. Existing datasets label such commits by tracing a fix back through the version history, which often points to the wrong commit. We therefore audit every case by hand against an explicit four-criterion definition of a VIC, so the benchmark does not inherit that label noise. Each case is annotated with gold context, 464 code blocks in total, each tagged by its role in the evidence for the vulnerability. We evaluate seven frontier models with precision, recall and F1 at three granularities (file, block, and line), scored separately on the context an agent viewed while exploring and on the context it finally declared as evidence. The gap between the two is the main finding. Every model opens most of the gold context while exploring, but reports only part of it as evidence. At the level of code blocks, the share of the gold context a model reports is 37 to 73 percentage points below the share it viewed. Qwen3-Coder-Next views 86.3% of the lines in annotated code blocks but cites only 12.9% in its final report. GPT-5.5, which cites the most, views 73% and reports 36%. These results highlight a gap between finding relevant code and selecting it for the final report, which verdict-level benchmarks cannot reveal.