评估NIST漏洞框架作为CWE的继任者在自动化漏洞分类中的表现
Evaluating the NIST Bugs Framework Against CWE as a Successor for Automated Vulnerability Classification
- The University of Alabama(阿拉巴马大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究评估NIST漏洞框架(BF)作为CWE继任者的可行性,通过专家标注和LLM自动化实验,证明BF在结构化程度和自动化友好性上优于CWE,但存在属性指导不足等空白。
AI中文摘要:
基于根本原因弱点的漏洞分类对于众多网络安全活动至关重要,其中通用弱点枚举(CWE)作为此类缺陷的公共存储库。然而,其重叠的条目导致了非正交结构,结果是同一漏洞被映射到多个弱点,使根本原因分析(RCA)和分类变得复杂。为解决此问题,NIST特别出版物800-231引入了漏洞框架(BF),该框架将漏洞组织为<原因,操作,后果>三元组,并将此类三元组链接成因果链,使漏洞携带其根本原因和汇聚点,而非单一的终端标签。然而,迄今为止,BF已被规范但尚未评估其在应对自动化分类挑战方面的性能,采用所需的证据尚未经过实证调查。我们使用一个系统筛选的、与CWE研究相关联的自动化通用漏洞与暴露(CVE)语料库,评估BF作为分类目标和CWE的补充。我们通过两项评估检验CVE到BF分类的可复现性。第一项是定性的:一项匿名化的评估者间研究,其中2名领域专家(SMEs)独立将13个CVE映射到四个BF轴上。标注者在原因和操作轴上表现出强一致性,而属性轴显示中等一致性。我们还测试了我们的自动化框架,在两种大型语言模型(LLM)部署下,使用不同预算进行可复现性分析。尽管存在局限性,如证据可用性和闭源软件缺乏可检索的修复提交,我们的发现支持BF比CWE更结构化且更易于自动化的主张。我们的探索揭示了BF的具体空白,包括属性指导不足。
英文摘要:
Vulnerability classification based on root cause weaknesses is essential for numerous cybersecurity activities, where the Common Weakness Enumeration (CWE) serves as a public repository of such flaws. However, its overlapping entries create a non-orthogonal structure. The result is the same vulnerability being mapped to multiple weaknesses, complicating Root Cause Analysis (RCA) and triage. To address this, NIST Special Publication 800-231 introduces the Bugs Framework (BF), which organizes vulnerabilities into <cause, operation, consequence> triples and links such triples into a causal chain, so that a vulnerability carries its root cause and its sink together instead of a single terminal label. To date, however, BF has been specified but not evaluated regarding its performance against the challenges to automated classification. The evidence required for adoption has not been investigated empirically. We evaluate BF as a classification target and a complement to CWE using a systematically screened corpus of automated Common Vulnerabilities and Exposures (CVEs) linked to CWE research. We assess the reproducibility of CVE-to-BF classification through two evaluations. The first is qualitative: an anonymized inter-rater study in which 2 subject-matter experts (SMEs) independently mapped 13 CVEs onto the four BF axes. Annotators showed strong agreement on the cause and operation axes, while the attribute axis indicated fair agreement. We also tested our automated framework across two large language model (LLM) deployments under different budgets for reproducibility analysis. Despite limitations, such as evidence availability and the absence of retrievable fix commits for closed-source software, our findings support the claim that BF is a more structured and automation-friendly framework than CWE. Our exploration reveals specific gaps in BF, including under-specified guidance on attributes.