用于从网络威胁情报中提取攻击技术的检索约束策略优化
Retrieval-Constrained Policy Optimization for Attack Technique Extraction from Cyber Threat Intelligence
浏览论文内容
中文总结 AI 辅助
研究针对CTI文本提取MITRE ATT&CK技术的挑战,提出两阶段框架TTP-R1,在四个CTI基准测试中F1值最优,子技术级F1超带检索增强的Claude Sonnet 4.5且推理速度更快
中文摘要 AI 辅助
将网络威胁情报(CTI)文本映射到MITRE ATT&CK攻击技术是结构化威胁分析的关键,但手动标注成本高昂且无法扩展。ATT&CK分类体系包含数百种攻击技术,单条CTI文本可能描述多种技术,导致准确完整的提取颇具挑战性。现有自动化方法存在不同不足:多标签分类器难以应对严重的类别不平衡和庞大的标签空间;而基于大语言模型(LLM)的方法——检索管道和微调生成器——优化的是标记级目标,将技术标注视为序列生成而非集合预测,缺乏对预测技术集是否正确完整的直接监督。我们提出TTP-R1,这是一个结合检索增强的监督微调(SFT)与使用可验证奖励的强化学习(RLVR)的两阶段框架。混合检索器首先将庞大的标签空间缩小为候选集,微调后的LLM学习选择正确的技术。随后,我们应用分组相对策略优化,其分解后的奖励直接监督预测技术集的精确率、召回率和输出格式。在四个CTI基准测试中,TTP-R1取得了最佳平均F1值,其子技术级F1较带检索增强的Claude Sonnet 4.5提升了7.4个百分点,且作为80亿参数模型在单GPU上部署时运行速度快28倍。
英文摘要
Mapping cyber threat intelligence (CTI) text to MITRE ATT&CK techniques is essential for structured threat analysis, yet manual annotation is costly and does not scale. The ATT&CK taxonomy comprises several hundred attack techniques, and a single CTI passage may describe multiple techniques, making accurate and complete extraction challenging. Existing automated approaches fall short in different ways: multi-label classifiers struggle with severe class imbalance and the large label space, while LLM-based methods--retrieval pipelines and fine-tuned generators--optimize token-level objectives that treat technique annotation as sequence generation rather than set prediction, lacking direct supervision on whether the predicted technique set is correct and complete. We propose TTP-R1, a two-stage framework that combines retrieval-augmented supervised fine-tuning (SFT) with reinforcement learning using verifiable rewards (RLVR). A hybrid retriever first narrows the large label space to a candidate set, and a fine-tuned LLM learns to select the correct techniques. We then apply Group Relative Policy Optimization with a decomposed reward that directly supervises the precision, recall, and output format of the predicted technique set. Across four CTI benchmarks, TTP-R1 achieves the best average F1, improving sub-technique-level F1 by 7.4 percentage points over Claude Sonnet 4.5 with retrieval augmentation, while running 28x faster when served as an 8B-parameter model on a single GPU.
发表机构
- Amazon Web Services(亚马逊云科技)
- Rutgers University(罗格斯大学)
机构由 AI 辅助整理,请以论文原文为准。