AI 中文总结
GraftyVul是通过将真实世界漏洞植入开源项目构建不安全程序的系统,生成212个跨5种语言、23个CWE类别的可验证漏洞程序,在多样性等指标上优于13个常用数据集,可用于评估漏洞修复系统。
AI 中文摘要
漏洞数据集是漏洞检测、自动修复、安全代码生成等各类安全研究的基础支撑,但现有数据集至少缺失以下三项理想属性中的一项:多样性(语言或漏洞类型)、可复现性/可执行性、真实性。为此,我们提出GraftyVul系统,该系统通过将真实世界漏洞植入开源项目来构建不安全程序。该系统使数据集基于真实场景中观察到的漏洞,同时利用已知的良好构建和测试环境,使漏洞验证脚本能够保证引入的漏洞成功改变程序行为。使用GraftyVul,我们生成了212个经过验证且可利用的不安全程序,涵盖5种编程语言(Python、TypeScript、Java、Go和C#),涉及23个CWE类别。为评估保真度,我们引入了一种与语言和上下文无关的语义嵌入方法,该方法通过漏洞的 sink(漏洞入口)、机制和宿主特征而非表面代码来比较漏洞。该方法在跨语言克隆和CWE分类任务上的表现优于标准代码嵌入。这些嵌入结果表明,GraftyVul生成的样本与其源漏洞保留了强语义特征。我们还将GraftyVul与13个广泛使用的数据集进行比较,发现它在实现有竞争力的多样性的同时,是唯一具有广泛语言和CWE覆盖范围的可复现漏洞利用数据集。最后,我们通过一项评估生产漏洞修复系统的工业案例研究,说明了GraftyVul的实际应用价值。
英文摘要
Vulnerability datasets underpin a wide range of security research, including vulnerability detection, automated remediation, and secure code generation. However, existing datasets sacrifice at least one of three desirable properties: diversity (of language or vulnerability type), reproducibility/executability, or realism. We therefore present GraftyVul, a system that constructs vulnerable programs by grafting real-world vulnerabilities into open-source projects. This grounds the dataset in vulnerabilities observed in real-world contexts while harnessing known good build and test environments, enabling exploit-verification scripts to guarantee that an introduced vulnerability successfully alters a program's behaviour. Using GraftyVul, we generate 212 verified and exploitable vulnerable programs spanning five programming languages (Python, TypeScript, Java, Go, and C#) across 23 CWE categories. To evaluate fidelity, we introduce a language- and context-agnostic semantic embedding that compares vulnerabilities by sink, mechanism and host-feature rather than surface code. This approach outperforms standard code embeddings on cross-language clone and CWE classification. These embeddings demonstrate that GraftyVul samples retain a strong semantic signature to their source vulnerability. We additionally compare GraftyVul against 13 widely used datasets, where it attains competitive diversity while being the only reproducible-exploit dataset with broad language and CWE coverage. Finally, we illustrate GraftyVul's practical utility through an industrial case study evaluating a production vulnerability remediation system.
Comments13 pages, 8 figures