AI 中文总结
研究针对CTI报告非结构化致实体和行为提取分析复杂的问题,构建含150份英文报告的数据集,以STIX 2.1图表示,含多种实体和关系且映射到MITRE ATT&CK。经评估建立参考数据集并评估开源LLMs,Qwen3.6:27B性能最佳,为相关研究提供基准。
AI 中文摘要
网络威胁情报(CTI)报告通常采用非结构化格式,这使得重要实体和对抗行为的提取与分析变得复杂。现有CTI研究虽提供了提取工具、知识图谱框架和MITRE ATT&CK映射数据集,但保留复杂实体关系和标准化对抗行为的精选报告级数据集仍然有限。本研究提出了一个由150份英文CTI报告组成的人工构建数据集,每个报告以基于STIX 2.1的图表示,包含4777个STIX实体、5817个STIX关系,以及1273个映射到269个独特MITRE ATT&CK企业技术和子技术的STIX攻击模式实体(对抗行为)。两名网络安全研究人员对25份随机抽样报告进行独立评估,结果显示评分者间有较高一致性。随后对分歧进行裁决以建立黄金标准参考数据集。对四个本地部署的开源大语言模型(LLMs)作为自动评判工具进行评估,Qwen3.6:27B总体性能最强,kappa分数最高达0.803,微F1分数超过92%,误报率低于5%。该数据集为CTI信息提取、知识图谱构建、事件分析和威胁归因提供了基准。研究结果还表明,本地部署的LLMs可支持人工评审员识别注释不一致问题,但专家验证仍然至关重要。
英文摘要
Cyber threat intelligence (CTI) reports are typically written in unstructured formats, which complicates the extraction and analysis of important entities and adversarial behaviors. Although existing CTI research provides extraction tools, knowledge-graph frameworks, and MITRE ATT&CK mapped datasets, curated report-level datasets that preserve complex entity relationships and normalized adversarial behaviors remain limited. To address this limitation, this study presents a manually constructed dataset of 150 English-language CTI reports, each represented as STIX 2.1 based graphs, which includes 4,777 STIX entities, 5,817 STIX relationships in total, and 1,273 STIX attack-pattern entities (adversarial behaviors) mapped to 269 unique MITRE ATT&CK Enterprise techniques and sub-techniques. Twenty five randomly sampled reports were independently assessed by two cybersecurity researchers, which shows substantial inter-rater agreement. Disagreements were subsequently adjudicated to establish a gold-standard reference dataset. Four locally deployed open-source LLMs were evaluated as automated judges against this adjudicated reference sample. Qwen3.6:27B achieved the strongest overall performance, with a maximum kappa score of 0.803, micro-F1 scores exceeding 92%, and false-positive rates below 5%. The dataset provides a benchmark for CTI information extraction, knowledge-graph construction, incident analysis, and threat attribution. The findings further indicate that locally deployed LLMs can support human reviewers in identifying annotation inconsistencies, but expert validation remains essential.
Comments6 pages, 2 figures, 9 tables