arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于证据检索的CTI报告调查线索生成

Evidence-Grounded Retrieval for Investigation Hunt Lead Generation from CTI Reports

Akash Prakash, Boubakr Nour, Makan Pourzandi, Chadi Assi, Mourad Debbabi

arXiv 2609.08790首次发表:更新:

发表机构

Concordia University; Ericsson Security Research(康考迪亚大学; 爱立信安全研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对威胁报告人工生成调查线索繁琐且难以扩展的问题,提出AHLERT系统,通过混合检索与本体接地生成环境感知的可操作线索,将平均F1提升约2倍并取得最高效能。

AI 中文摘要

威胁狩猎日益依赖于将非结构化知识(如网络威胁情报报告)转化为可操作的调查线索:即基于可观察工件和对手技术的简明、可调查的假设。手动生成此类线索是一项繁琐且难以扩展的任务。现有的自动化方法止步于实体层,忽略防御者的运营环境,并孤立地分析每份报告。为解决这些不足,我们提出了AHLERT系统,该系统通过以下方式自动从威胁报告中提取相关、环境感知的调查线索:(i)一种混合检索器,将密集向量搜索与基于MITRE ATT&CK种子的知识图谱上的多跳遍历相结合;(ii)一种本体接地检索增强生成方法,将每条线索约束在防御者自身的资产和控制范围内;(iii)一个与LLM无关的框架,输出结构化、可直接操作的线索,而非松散的失陷指标。我们在针对知名APT的公开CTI报告上,使用多种专有和开源权重模型评估了AHLERT。混合证据检索与本体接地相比单路平坦RAG基线,将平均F1值提高了约2倍(从0.44提升至0.85),且AHLERT相比现成的LLM模型取得了最高效能分数(约86.95%)。

英文摘要

Threat hunting increasingly depends on converting unstructured knowledge (e.g., Cyber Threat Intelligence reports) into actionable hunt leads: concise, investigable hypotheses grounded in observable artifacts and adversary techniques. Producing such leads manually is a tedious and hard-to-scale task. Existing automated approaches stop at the entity layer, ignore the defender's operational environment, and analyze each report in isolation. To address these gaps, we introduce AHLERT, a system that automatically extracts relevant, environment-aware, and hunt leads from threat reports through (i) a hybrid retriever that combines dense vector search with multi-hop traversal over a knowledge graph seeded with MITRE ATT&CK; (ii) an ontology-grounding retrieval-augmented generation method that constrains each lead to the defender's own assets and controls; and (iii) an LLM-agnostic framework that emits structured, directly actionable leads rather than loose indicators of compromise. We evaluate AHLERT on public CTI reports for well-known APTs across multiple proprietary and open-weight models. Hybrid evidence retrieval with ontology grounding raises mean F1 by ~2x (0.44 to 0.85) over a single-route flat-RAG baseline, and AHLERT attains the highest effectiveness score (~86.95%) compared with off-the-shelf LLM models.

CommentsAccepted for presentation and publication at the 2026 IEEE Conference on Dependable and Secure Computing (DSC) - Workshop: Cyber Resilience & Attack Intelligence (CRAI)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑