AI 中文总结
研究从自由文本漏洞描述映射CVE到MITRE ATT&CK技术,训练多标签分类器提升召回率等指标。探讨大语言模型辅助标签扩展,发现其受评估噪声影响,无法可靠改进,分类器受标签质量而非数据集大小限制。
AI 中文摘要
我们提出了一个可重现的管道,用于从自由文本漏洞描述中将通用漏洞披露(CVE)映射到MITRE ATT&CK企业技术。我们没有依赖CWE->CAPEC->ATT&CK推导链(我们对其表扩展工件进行了量化),而是在一个由来自MITRE威胁情报防御中心专家映射的1207个CVE组成的精心策划的黄金数据集上训练了一个多标签分类器。与零样本嵌入相似性基线相比,所得模型的召回率@5提高了约一倍,并且改善了每个排名指标。然后,我们研究了大语言模型辅助标签是否可以扩展黄金数据集。初步实验得出了相互矛盾的结论:单次运行表明性能下降,而对五个随机种子进行平均则表明有小幅提升。然而,一项独立复制和扩展规模研究(额外增加100-984个CVE)表明,明显改善是一种评估工件。大语言模型生成的标签与专家注释的一致性约为0.39在任何扩展规模下都没有提供可靠的改进,并且在增加约1000个CVE时会降低稀有技术覆盖率(宏F1下降0.04)。根本原因是评估噪声。在小测试分割上选择检查点有效地在许多噪声评估中实现了最大化,在其他条件相同的运行之间产生了高达0.05的召回率@5差异。使用基于验证分割检查点选择的校正协议,仅黄金模型的召回率@5为(0.673±0.019),重复决定性实验证实了大语言模型扩展的无效结果。最终的规模研究表明,额外的专家策划数据持续提高性能,而大语言模型标记的数据则不然,这表明分类器受标签质量而非数据集大小的限制。所有数据集、模型、代码和训练日志均已公开发布。
英文摘要
We present a reproducible pipeline for mapping Common Vulnerabilities and Exposures (CVEs) to MITRE ATT&CK Enterprise techniques from free-text vulnerability descriptions. Rather than relying on the CWE->CAPEC->ATT&CK derivation chain, whose table-expansion artifacts we quantify, we train a multi-label classifier on a curated gold dataset of 1,207 CVEs from expert MITRE Center for Threat-Informed Defense mappings. The resulting model approximately doubles recall@5 compared with a zero-shot embedding-similarity baseline and improves every ranking metric. We then investigate whether LLM-assisted labeling can extend the gold dataset. Initial experiments suggest contradictory conclusions: a single run indicates degraded performance, while averaging over five random seeds suggests a small gain. However, an independent replication and an expansion-size study (100--984 additional CVEs) show that the apparent improvement is an evaluation artifact. LLM-generated labels, with approximately 0.39 agreement with expert annotations, provide no reliable improvement at any expansion size and reduce rare-technique coverage at around 1,000 added CVEs (macro-F1 decreases by 0.04). The root cause is evaluation noise. Selecting checkpoints on a small test split effectively maximizes over many noisy evaluations, producing recall@5 differences of up to 0.05 between otherwise identical runs. Using a corrected protocol based on validation-split checkpoint selection, the gold-only model achieves recall@5 of (0.673 \pm 0.019), and repeating the decisive experiment confirms the null result for LLM expansion. A final scaling study shows that additional expert-curated data consistently improves performance, whereas LLM-labeled data does not, indicating that the classifier is limited by label quality rather than dataset size. All datasets, models, code, and training logs are publicly released.