推理拓扑至关重要:基于LLM的网络安全分析对照研究
Reasoning Topology Matters: A Controlled Study of LLM-Based Cybersecurity Analysis
- University of Turku(图尔库大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出安全推理拓扑,通过对照实验证明图推理结构在网络安全分析中显著优于线性、分支及少样本提示,准确率提升9.8-12.2个百分点,且效果跨模型一致。
AI中文摘要:
大型语言模型(LLMs)在网络安全领域中的应用日益广泛,而准确的分析通常需要对复杂且异构的数据进行多步骤、依赖上下文的推理。然而,现有的提示方法通常侧重于引导推理,而未明确考虑中间推理步骤在结构上如何组织。我们引入了安全推理拓扑(Security Reasoning Topology),通过三种代表性结构对推理进行建模:线性(Linear)、分支(Branching)和图(Graph)。为评估其效果,我们在三个网络安全数据集上进行了对照实验,涵盖MITRE ATT&CK网络流量、网络威胁情报(CTI)和CVE漏洞分析。我们评估了多个LLM,包括Llama 2(7B、13B、70B)、GPT-5.1和Mistral Large 3,同时保持任务输入一致,并通过系统级提示控制推理结构。结果表明,推理拓扑对性能有显著影响:图推理实现了最高的整体准确率,在多个数据集上比少样本提示提高了9.8至12.2个百分点,而分支推理提供了强有力的中间解决方案。结果进一步表明,推理拓扑的影响在不同模型家族和规模中保持一致,凸显了推理拓扑作为基于LLM的网络安全分析中的一个重要设计因素。
英文摘要:
Large Language Models (LLMs) are increasingly used in cybersecurity, where accurate analysis often requires multi-step and context-dependent reasoning over complex and heterogeneous data. However, existing prompting approaches typically focus on eliciting reasoning without explicitly considering how intermediate reasoning steps are structurally organized. We introduce Security Reasoning Topology, which models reasoning through three representative structures: Linear, Branching, and Graph. To evaluate their effects, we conduct controlled experiments on three cybersecurity datasets covering MITRE ATT&CK network traffic, cyber threat intelligence (CTI), and CVE vulnerability analysis. We evaluate multiple LLMs, including Llama 2 (7B, 13B, 70B), GPT-5.1, and Mistral Large 3, while keeping task inputs consistent and controlling reasoning structure through system-level prompting. Results show that reasoning topology substantially affects performance: Graph reasoning achieves the highest overall accuracy, improving over few-shot prompting by 9.8-12.2 percentage points across datasets, while Branching provides a strong intermediate solution. The results further show that the effect of reasoning topology remains consistent across model families and scales, highlighting reasoning topology as an important design factor for LLM-based cybersecurity analysis.