CTIFoundry:用于网络威胁情报的智能体原生语料库支架
CTIFoundry: An Agent-Native Corpus Scaffold for Cyber Threat Intelligence
查看机构详情
- Virginia Tech(弗吉尼亚理工大学)
- Amazon(亚马逊公司)
- Northwestern University(西北大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
CTIFoundry是用于网络威胁情报的智能体原生语料库支架,在CTIConnect基准测试中可使智能体F1提升0.19至0.28,小型模型在其上表现优于旗舰模型,且无需增加搜索工作量。
中文摘要 AI 辅助
网络威胁情报(CTI)越来越多地被大型语言模型(LLM)智能体而非人类分析师使用,这些智能体会在查询时构建多步骤调查。这一转变的利用侧已迅速成熟(规划循环、工具协议、上下文管理),但语料库侧尚未成熟:威胁报告和漏洞数据库仍以检索增强生成的形式打包,作为嵌入索引后的不透明块存在。我们认为,这一底层架构而非模型能力是智能体CTI调查的瓶颈,并提出了CTIFoundry,一种智能体原生语料库支架。在构建时,CTIFoundry实现了CTI语料库的潜在结构:一个基于四个权威知识库(CVE、CWE、CAPEC、ATT&CK)的确定性本体图,其官方交叉引用成为类型化、可遍历的边;一个基于跨度的报告层,其规范的、别名解析的跨供应商实体索引带有来源的块;以及混合密集+词汇检索表面。在查询时,该结构通过安装在标准开源智能体 harness 上的七个类型化工具和三个程序技能暴露出来。在公开的CTIConnect基准测试中,仅交换动作表面就使相同 harness 的智能体在四个模型、两个提供商的面板上的整体F1提升了+0.19至+0.28:CTIFoundry上的小型模型超过了平面底层架构上的旗舰模型,且这种提升并非以搜索工作量为代价,因为在两个Claude模型上,支架智能体在大约一半的工具调用中更准确。消融分析将此归因于:类型化结构贡献了更大的份额,程序技能将结构转化为规范性,两者以超加性方式组合,因为技能仅绑定到存在的结构。
英文摘要
Cyber threat intelligence (CTI) is increasingly consumed not by human analysts but by LLM agents that compose multi-step investigations at query time. The harness side of this shift has matured rapidly, but the corpus side has not: threat reports and vulnerability databases are still packaged for retrieval-augmented generation, as opaque chunks behind an embedding index. We argue that this substrate, not model capability, is the bottleneck on agentic CTI investigation, and present CTIFoundry, an agent-native corpus scaffold. At build time, CTIFoundry materializes the latent structure of a CTI corpus: a deterministic ontology graph over four authoritative knowledge bases (CVE, CWE, CAPEC, ATT&CK) whose official cross-references become typed, traversable edges; a span-grounded report layer whose canonical, alias-resolved cross-vendor entities index provenance-carrying chunks; and hybrid dense+lexical retrieval surfaces. At query time this structure is exposed through seven typed tools and three procedural skills mounted on a stock, widely-used open-source agent harness. On the public CTIConnect benchmark, swapping only the action surface lifts the identically-harnessed agent from 0.610 to 0.829 overall F1 with gpt-5.4 and from 0.470 to 0.745 with claude-haiku-4-5: a small model on CTIFoundry surpasses a flagship model on the flat substrate. The scaffolded agent is simultaneously more accurate and more efficient: on both Claude models it answers with roughly half the tool calls per question. The ablation distills design principles for matching corpus scaffolding to data modality, in CTI and beyond. Build-time validation guarantees zero fabricated identifiers by construction, and the scaffold sustains 1,168 investigations end-to-end at about 2.6 cents each.