MITRE-SAGE:一种多智能体网络安全问答模型
MITRE-SAGE: A Multi-Agent Cybersecurity Question-Answering Model
浏览论文内容
中文总结 AI 辅助
针对网络安全领域LLM的不足,研究提出多智能体框架MITRE-SAGE,结合MITRE-QA基准验证其在多项网络安全问答任务中优于基线方法,轻量级配置表现突出。
中文摘要 AI 辅助
有效的网络安全操作需要及时准确地分析大规模异构安全信息;然而,分析师日益面临信息过载、警报疲劳和决策时间受限的问题。尽管大型语言模型(LLM)在问答(QA)方面展现出良好的能力,但它们在网络安全领域的有效性受到领域知识不足、存在幻觉倾向以及难以同时捕捉语义和结构关系的限制。本研究提出了MITRE-SAGE,一种多智能体检索增强生成框架,它整合了语义和结构网络安全知识,以提升基于LLM的问答系统的可靠性和可解释性。通过将复杂任务分解为查询解析、证据检索和答案合成,MITRE-SAGE有效支持漏洞评估、威胁画像和关系提取等网络安全任务。此外,我们提出了MITRE-QA,一个包含3000个问答对的综合基准,用于在多样化的网络安全知识任务中评估LLM,并使用该基准系统地评估MITRE-SAGE与代表性基线方法的性能。大量实验表明,MITRE-SAGE始终优于独立LLM和传统RAG方法。值得注意的是,由Qwen2.5-7B子智能体和Qwen2.5-14B协调器组成的轻量级配置在8项基准任务中的5项上实现了更优性能,表明所提出的多智能体框架的有效性。这些结果凸显了MITRE-SAGE作为一种可扩展且可解释的可靠网络安全问答方法的潜力,而MITRE-QA则为未来研究提供了标准化基准。
英文摘要
Effective cybersecurity operations require timely and accurate analysis of large-scale heterogeneous security information; however, analysts increasingly struggle with information overload, alert fatigue, and time-constrained decision-making. Although large language models (LLMs) have demonstrated promising capabilities for question answering (QA), their effectiveness in cybersecurity remains limited by insufficient domain knowledge, a tendency to hallucinate, and difficulties in capturing both semantic and structural relationships. This work proposes MITRE-SAGE, a multi-agent retrieval-augmented generation framework that integrates semantic and structural cybersecurity knowledge to improve the reliability and interpretability of LLM-based QA systems. By decomposing complex tasks into query interpretation, evidence retrieval, and answer synthesis, MITRE-SAGE effectively supports cybersecurity tasks such as vulnerability assessment, threat profiling, and relationship extraction. Furthermore, we propose MITRE-QA, a comprehensive benchmark comprising 3,000 question-answer pairs for evaluating LLMs across diverse cybersecurity knowledge tasks, and use it to systematically evaluate MITRE-SAGE against representative baseline methods. Extensive experiments demonstrate that MITRE-SAGE consistently outperforms standalone LLMs and conventional RAG approaches. Notably, a lightweight configuration comprising Qwen2.5-7B sub-agents and a Qwen2.5-14B orchestrator achieves superior performance on five of the eight benchmark tasks, indicating the effectiveness of the proposed multi-agent framework. The results highlight the potential of MITRE-SAGE as a scalable and interpretable approach for reliable cybersecurity QA, while MITRE-QA provides a standardized benchmark for future research.
发表机构
- University of Guilan(吉兰大学)
机构由 AI 辅助整理,请以论文原文为准。