发表机构
Guangdong Provincial Key Laboratory of Novel Security Intelligence Technologies; School of Computer Science and Technology; Harbin Institute of Technology; New Network Research Department; Peng Cheng Laboratory; Great Bay University(广东新型安全情报技术重点实验室; 计算机科学与技术学院; 哈尔滨工业大学; 新网络研究部; 鹏城实验室; 大湾区大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对CTI报告非结构化无法用于自动化攻击路径推理的问题,提出将攻击步骤建模为含条件和行为的单元,经多阶段管道结合大语言模型提取并规范化,编译成规则推理,在数据集上有更好表现,能有效提取可达攻击链。
AI 中文摘要
网络威胁情报(CTI)报告详细描述了现实世界中的攻击过程,但其非结构化叙述无法直接用于自动化攻击路径推理。现有CTI提取方法未对攻击步骤的执行条件和结果状态建模,无法支持跨多阶段攻击链的状态匹配和可达性分析。本文提出一个自动化框架,将每个攻击步骤建模为包含前置条件、攻击行为和后置条件的攻击单元来提取可达攻击链。通过多阶段管道结合大语言模型提取攻击行为框架、恢复条件并规范化,编译成Datalog规则进行攻击目标可达性推理。在含334个经人工验证标注步骤的数据集中,该框架在恢复攻击行为方面有更高覆盖率,生成的攻击单元更完整一致。Datalog推理在20个报告中的19个达到指定攻击目标,反向搜索产生34条攻击路径。
英文摘要
Cyber Threat Intelligence (CTI) reports richly describe real-world attack processes, but their unstructured narratives cannot be directly used for automated attack-path reasoning. Existing CTI extraction methods focus on indicators, entities, or TTP labels without modeling the execution conditions and resulting states of each attack step, so the extracted knowledge supports neither state matching nor reachability analysis across multi-stage attack chains. This paper proposes an automated framework that extracts reachable attack chains by modeling each attack step as an attack unit of preconditions, an attack behavior, and postconditions. A multi-stage pipeline assisted by large language models (LLMs) extracts attack behavior skeletons, recovers their preconditions and postconditions, normalizes them into predefined predicates, and repairs broken dependencies; the resulting units are compiled into Datalog-style rules for attack-goal reachability reasoning. On a dataset of 20 CTI reports containing 334 human-validated annotated steps, our framework achieves higher annotated-step coverage than representative CTI extraction systems in recovering attack behaviors. Moreover, by explicitly generating preconditions and postconditions, it produces attack units that are more complete and consistent than those generated by end-to-end LLM baselines. On the extracted chains, Datalog inference reaches the specified attack goal in 19 of 20 reports, while backward search yields 34 attack paths under the generated rules. The source code and experimental artifacts are available in an anonymized repository. .
Comments16 pages, 3 figures