AI 中文总结
该研究构建基于Seclometry的CWEAgent,对NVD中15556个开源CVE的CWE标签进行大规模审计,发现仅49.70%标签准确,存在结构性标注错误问题。
AI 中文摘要
国家漏洞数据库(NVD)中的CWE标签被广泛用作漏洞搜索、扫描器评估、基准构建、基于学习的安全工具及漏洞优先级排序的基准。然而,尽管人们对NVD的补充积压问题日益担忧,且存在不准确、模糊或缺失标签的传闻报告,其可靠性尚未得到大规模系统测量。本文提出一种基于代码语义的大规模NVD中CWE标注质量测量方法。我们构建了CWEAgent,这是一种经过验证的审计工具,基于Seclometry(一种结构化的漏洞语义表示,可捕获漏洞代码的根本原因、触发条件、被违反的安全属性、利用机制及影响)。在一个由100个开源CVE组成的人工策划基准上,CWEAgent达到85%的Top-1准确率和92%的歧义感知准确率。将CWEAgent应用于2017-2026年披露的15556个开源CVE,我们发现仅有49.70%的NVD CWE标签与基于代码的标签完全匹配;另有31.37%是分类歧义下的可辩护替代标签,而3.63%是与证据不一致的可能错误。标签可靠性随分配机构和弱点类型的不同而显著变化,且明显的项目级或语言级差异在很大程度上是这些潜在弱点类型的构成效应。与证据不一致的标签也随时间推移而增加。通过对434个已确认错误标签的人工审查,我们识别出六种反复出现的错误模式,表明CWE噪声是漏洞元数据中的结构性问题,而非孤立的标注错误。
英文摘要
CWE labels in the National Vulnerability Database (NVD) are widely treated as ground truth for vulnerability search, scanner evaluation, benchmark construction, learning-based security tools, and vulnerability prioritization. Yet their reliability has not been systematically measured at scale, despite growing concerns about NVD's enrichment backlog and anecdotal reports of inaccurate, ambiguous, or missing labels. This paper presents a large-scale, code-semantics-grounded measurement of CWE labeling quality in NVD. We build CWEAgent, a validated auditing instrument based on seclometry, a structured representation of vulnerability semantics that captures the root cause, trigger condition, violated security property, exploit mechanism, and impact of vulnerable code. On a manually curated benchmark of 100 open-source CVEs, CWEAgent achieves 85% top-1 accuracy and 92% ambiguity-aware accuracy. Applying CWEAgent to 15,556 open-source CVEs disclosed from 2017-2026, we find that only 49.70% of NVD CWE labels exactly match the code-grounded label. Another 31.37% are defensible alternatives under taxonomy ambiguity, while 3.63% are evidence-inconsistent likely errors. Label reliability varies sharply by assigning organization and weakness type, and apparent project- or language-level differences are largely composition effects of those underlying weakness types. Evidence-inconsistent labels have also increased over time. Through manual review of 434 confirmed mislabels, we identify six recurring error patterns, showing that CWE noise is a structural problem in vulnerability metadata rather than isolated annotation mistakes.