arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过检索增强生成和小语言模型融合网络威胁情报以实现丰富的威胁表示

Merging Cyber Threat Intelligence Through Retrieval-Augmented Generation and Small Language Models for Rich Threat Representation

Nicola Deidda, Leonardo Regano, Alessandro Sanna, Davide Maiorca, Giorgio Giacinto

arXiv 2609.07280首次发表:更新:

发表机构

University of Cagliari; IMT School for Advanced Studies Lucca; Consorzio Interuniversitario Nazionale per l’Informatica(卡利亚里大学; 卢卡高等研究学院; 意大利国家大学计算机联盟)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出一种结合检索增强生成与小语言模型的自动化流水线,从异构CTI来源生成增强攻击图,以支持预防、检测和响应,并在10个真实案例中验证其有效性。

AI 中文摘要

现代网络安全运营依赖于从异构来源收集的网络威胁情报(CTI),这些来源包括半结构化的威胁表示、失陷指标(IoC)和叙事性技术报告。然而,这些工件在孤立状态下往往不足以重建攻击如何展开、每个步骤在何种条件下可行以及攻击留下哪些痕迹。在实践中,分析师必须手动关联分散在多个且仅部分结构化来源中的部分证据,这延迟了有效预防、检测和响应措施的设计。为解决这一差距,我们提出了一种自动化流水线,从异构CTI来源中推导出可操作的网络攻击表示。该流水线结合了检索增强生成(RAG)架构与可本地部署的小语言模型(SLM),用于整合此类证据并推断缺失的操作细节。从半结构化威胁表示和辅助CTI文档出发,该流水线生成一个增强的攻击图,该图捕获攻击的粗略、与战术对齐的进展,并为每个步骤标注显式的前置条件和后置条件以及增强描述。这种表示通过暴露执行要求来支持预防,通过突出可观察痕迹来支持检测,通过阐明攻击的时间进展来支持响应。然后,由于缺乏具有真实世界攻击时间演化地面真值信息的验证数据集,我们在10个涵盖多种威胁类型(包括通过钓鱼投递的后门和分阶段下载器)的真实世界案例研究上测试了完整流水线。对10个真实世界案例研究的手动评估提供了初步证据,表明生成的图与预期的攻击进展一致,表明所提出的方法可以通过将分散的CTI证据整合为结构化且可操作的攻击视图来支持分析师。

英文摘要

Modern cybersecurity operations rely on CTI collected from heterogeneous sources, including semi-structured threat representations, IoCs, and narrative technical reports. However, these artifacts are often insufficient in isolation to reconstruct how an attack unfolds, under which conditions each step is feasible, and which traces it leaves behind. In practice, analysts must manually correlate partial evidence scattered across multiple and only partially structured sources, delaying the design of effective prevention, detection, and response actions. To address this gap, we propose an automated pipeline that derives an actionable representation of a cyberattack from heterogeneous CTI sources. The pipeline combines a RAG architecture with a locally deployable SLM, used to consolidate such evidence and infer missing operational details. Starting from a semi-structured threat representation and auxiliary CTI documents, the pipeline produces an enriched Attack Graph that captures a coarse, tactic-aligned progression of the attack and annotates each step with explicit pre-conditions and post-conditions, and an enriched description. This representation supports prevention by exposing execution requirements, detection by highlighting observable traces, and response by clarifying the temporal progression of the attack. Then, due to the lack of validated datasets with ground-truth information on the temporal evolution of real-world attacks, we test the complete pipeline on 10 real-world case studies spanning multiple threat types, including backdoors and staged downloaders delivered via phishing. A manual assessment across 10 real-world case studies provides initial evidence that the generated graphs are consistent with expected attack progressions, indicating that the proposed approach can support analysts by consolidating dispersed CTI evidence into a structured and actionable view of attacks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑