基于结构化风险指标的零样本威胁检测大语言模型
LLMs for Zero-Shot Threat Detection via Structured Risk Indicators
浏览论文内容
中文总结 AI 辅助
提出两阶段LLM框架,结合RAG从异构安全日志零样本检测内部威胁与APT,在两个基准数据集上优于现有模型,风险指标质量是性能关键驱动因素
中文摘要 AI 辅助
我们提出了一种两阶段大语言模型(LLM)框架,用于从异构安全日志中零样本检测内部威胁和高级持续性威胁(APT)。该框架将用户活动建模为时间线,并结合检索增强生成(RAG)以提供来自每个用户历史活动的个性化行为上下文。该框架并非直接从原始日志执行端到端分类,而是首先生成结构化、可解释的特定威胁风险指标集,随后在时间序列上联合分类以捕捉跨越多个时间的攻击模式。我们在两个基准数据集上评估了该框架:用于内部威胁检测的CERT r5.2和用于APT检测的PicoDomain,在检索和非检索设置下,使用两个开源权重LLM的四种组合进行测试。所有配置均优于先前的最先进基于LLM的框架(GABM),最佳配置在CERT r5.2上将F1分数提高了11.40个百分点,在PicoDomain上提高了31.50个百分点。结果进一步表明,检索主要通过生成更具判别力的风险指标来使较弱的LLM受益,而较强的模型在无检索上下文的情况下可达到相当的性能。LLM在两个阶段的最有效分配取决于数据集。这些发现表明,生成的风险指标的质量是零样本网络威胁检测性能的主要驱动因素。
英文摘要
We propose a two-stage large language model (LLM) framework for zero-shot detection of insider threats and advanced persistent threats (APTs) from heterogeneous security logs. The framework models user activity as chronological timelines and incorporates retrieval-augmented generation (RAG) to provide personalised behavioural context from each user's historical activity. Rather than performing end-to-end classification directly from raw logs, it first generates structured, interpretable sets of threat-specific risk indicators, which are then classified jointly across temporal sequences to capture attack patterns spanning multiple windows.The framework is evaluated on two benchmark datasets, CERT r5.2 for insider threat detection and PicoDomain for APT detection, using four combinations of two open-weight LLMs under both retrieval and non-retrieval settings. All configurations outperform the previous state-of-the-art LLM-based framework (GABM), with the best configuration improving the F1-score by 11.40 percentage points on CERT r5.2 and 31.50 percentage points on PicoDomain. Results further show that retrieval mainly benefits weaker LLMs by generating more discriminative risk indicators, whereas stronger models achieve comparable performance without retrieved context. The most effective assignment of LLMs to the two stages depends on the dataset. These findings show that the quality of the generated risk indicators is the main driver of zero-shot cyber threat detection performance.
发表机构
- University of Queensland(昆士兰大学)
机构由 AI 辅助整理,请以论文原文为准。