软件漏洞分析中的数据问题:制品、质量与使用
The Data Problem in Software Vulnerability Analysis: Artifacts, Quality, and Consumption
浏览论文内容
中文总结 AI 辅助
本研究以1522篇论文及111篇锚定论文为基础,揭示软件漏洞分析数据集在真实性、标签证据等属性上的表现,发现代码样本数据集真实性最低,数据泄露等问题未被充分处理。
中文摘要 AI 辅助
基于学习和大语言模型(LLM)的软件漏洞分析的可信度仅与其训练和评估所用数据相当,但这些数据很少被视为一等对象进行审查。本研究通过以数据集为中心的分类法调查漏洞分析背后的数据,该分类法区分了制品是什么(代码、元数据、补丁、测试/概念验证(PoC)、推理、轨迹)、其质量如何(真实性、标签证据、规模、多样性、数据泄露、可用性)以及其用途。从系统汇编的2016-2026年及更早基础工作的1522篇论文语料库中,我们深度编码了111篇锚定论文的分层集合,为每个经 rubric 评分的肯定值提供逐字文本支撑,并按属性报告其被研究的程度以及数据集在该属性上的表现。结果勾勒出一条证据阶梯:可执行制品是唯一主要类型,其中24个数据集中有15个被评为同时符合真实世界标准且标签经过独立检查;而代码样本数据集——自动标记语料库和锚定集中最大的类别——真实性最低:41个数据集中有20个的漏洞来自真实项目或通用漏洞披露(CVE),但仅有3个将样本保留在代码部署的单元级别,且仅有2个同时满足这两点(不过这些是粗略的组件测试,仅有1个代码样本数据集符合代码手册更严格的全上下文真实世界等级)。其中,在适用数据泄露问题的90个数据集中,有49个未处理该问题;超过四分之一的数据集未提及可用性;推理数据仅在近期出现且大多为模型生成;而在对整个语料库进行筛查并对所有候选者进行全文检查后,主要轨迹语料库仅剩下3个数据集,其余数据集的轨迹则是在基准和测试 harness 之后次要发布的。
英文摘要
Learning- and LLM-based software vulnerability analysis is only as trustworthy as the data it is trained and evaluated on, yet that data is rarely examined as a first-class object. We investigate the data behind vulnerability analysis through a dataset-centric taxonomy that separates what an artifact is (code, metadata, patches, tests/PoCs, reasoning, traces), how good it is (realism, label evidence, scale, diversity, leakage, availability), and what it is used for. From a systematically assembled corpus of 1522 papers covering 2016-2026 plus foundational earlier work we deep-code a tiered set of 111 anchor papers, backing every affirmative rubric-graded value with a verbatim span, and we report, per attribute, both how much it has been studied and how well datasets achieve it. The results trace an evidence ladder: executable artifacts are the only major type where 15 of the 24 datasets are both graded real-world and carry labels that received an independent check, while code-sample datasets-the largest category in both the auto-tagged corpus and the anchor set-are the least realistic: 20 of the 41 draw their vulnerabilities from authentic projects or CVEs, but only 3 keep the sample at the unit the code is deployed in, and only 2 do both-though these are coarse component tests, and just one code-sample dataset meets the codebook's stricter full-context real-world grade. Among these, leakage goes unaddressed by 49 of the 90 datasets where it applies, more than a quarter say nothing about availability, reasoning data has arrived only recently and is mostly model-generated, and primary trace corpora remain limited to three datasets, the total after a corpus-wide screen and a full-text check of every candidate it surfaced, with further datasets releasing traces secondarily behind benchmarks and harnesses.
发表机构
- Oakland University(奥克兰大学)
- Macau University of Science and Technology(澳门科技大学)
- Kent State University(肯特州立大学)
- University at Buffalo, SUNY(纽约州立大学布法罗分校)
机构由 AI 辅助整理,请以论文原文为准。