VICBench:一个用于代码漏洞检测的多语言基准测试集
VICBench: A Multi-Language Benchmark for Code Vulnerability Detection
查看机构详情
- University of Pittsburgh(匹兹堡大学)
- Purdue University(普渡大学)
- Amazon Web Services(亚马逊网络服务)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究构建了多语言代码漏洞检测基准VICBench,包含100个CVE对应的VIC,规模显著大于现有数据集,经评估现有漏洞检测算法性能有限,该基准可用于可靠评估漏洞检测方法。
中文摘要 AI 辅助
评估安全漏洞检测工具需要包含漏洞引入提交(VICs)的基准数据集,VICs是首次将漏洞引入代码库的提交,对于确定易受攻击的软件版本范围至关重要。现有漏洞数据集存在编程语言覆盖有限、补丁复杂度受限、项目范围狭窄的问题。我们通过人类专家与智能体工作流的双重标注,创建了VICBench基准测试集,包含来自Python、Java、C++三类语言的88个项目中100个CVE对应的100个经核实的VIC,覆盖48种CWE类型。VICBench包含平均38.6行的复杂真实世界漏洞修复代码,以及对应平均252.5行的VIC,规模显著大于现有研究。我们的评估显示,最先进的算法V-SZZ和LLM4SZZ仅达到33.3%-40.1%的F1值,证实使用现有方法仍需大量手动工作,VICBench可实现对漏洞检测方法的可靠评估。
英文摘要
Evaluating security vulnerability detection tools requires benchmark datasets with vulnerability-inducing commits (VICs) - the commits that first introduce vulnerabilities into codebases. VICs are essential for determining the full range of vulnerable software versions. Existing vulnerability datasets suffer from limited programming language coverage, restricted patch complexity, and narrow project scope. Through our dual annotation by human experts and an agentic workflow, we create a benchmark - VICBench - of 100 verified VICs for 100 CVEs across 88 projects in Python, Java, and C++, covering 48 CWE types. VICBench features complex real-world vulnerability fixes averaging 38.6 lines and corresponding VICs of 252.5 lines - significantly larger than prior work. Our evaluation shows that state-of-the-art algorithms V-SZZ and LLM4SZZ achieve only 33.3%-40.1% F1, confirming that using existing approaches still entails significant manual effort. VICBench enables robust evaluation of vulnerability detection approaches.