SERA-IDS:基于结构化经验检索增强的小语言模型入侵检测
SERA-IDS: Structured Experience Retrieval-Augmented Intrusion Detection with Small Language Models
- Florida Polytechnic University(佛罗里达理工大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
SERA-IDS通过小语言模型从分类错误中生成结构化规则并检索匹配,显著提升入侵检测的宏F1,使小型本地模型性能超越现有方法。
AI中文摘要:
大型语言模型(LLM)为网络入侵检测提供了一种灵活的方法,但在没有流量特定决策边界的情况下,直接对数值流记录进行分类可能不可靠。检索增强生成(RAG)使LLM能够利用外部经验;然而,自由形式的文本经验要求模型推断区分流量类别的数值条件。我们提出SERA-IDS,一种从分类错误中学习决策知识的结构化经验检索增强入侵检测框架。一个小语言模型(SLM)分析误分类流,并生成包含行为描述、特征条件、类别混淆信息、工具派生置信度和来源的结构化规则。一个置信度门控将受支持的规则纳入经验库。在测试期间,经验库被冻结,检索到的类别多样化规则与查询流进行条件匹配,并作为证据提供给决策SLM。同一个经验库使用三个小型、可本地部署的SLM进行评估,即Llama 3.1、Phi-4:14B和Qwen2.5:7B,无需微调或付费托管推理。在测试集上,SERA-IDS将NF-BoT-IoT上的宏F1分别从11.66%、9.26%和7.83%提升至86.46%、69.50%和83.51%,在NF-ToN-IoT上从3.82%、6.17%和1.94%提升至84.16%、87.66%和83.37%。这些结果表明,结构化经验检索使小型、免费且可本地部署的SLM能够实现强大的入侵检测性能,包括超过我们比较中报告的现有工作的性能。
英文摘要:
Large language models (LLMs) offer a flexible approach to network intrusion detection, but direct classification of numerical flow records can be unreliable without traffic- specific decision boundaries. Retrieval-augmented generation (RAG) enables LLMs to leverage external experience; however, free-form textual experience requires the model to infer the numerical conditions that distinguish traffic classes. We propose SERA-IDS, a Structured Experience Retrieval-Augmented Intru- sion Detection framework that learns decision knowledge from classification errors. A Small Language Model (SLM) analyzes misclassified flows and generates structured rules containing behavioral descriptions, feature conditions, class-confusion infor- mation, tool-derived confidence, and provenance. A confidence gate admits supported rules into the Experience Library. During testing, the library is frozen, and retrieved class-diverse rules are condition-matched to the query flow and provided as evidence to the decision SLM. The same library is evaluated with three small, locally deployable SLMs, namely Llama 3.1, Phi-4:14B, and Qwen2.5:7B, without fine-tuning or paid hosted inference. On the test sets, SERA-IDS improves macro F1 on NF-BoT- IoT from 11.66%, 9.26%, and 7.83% to 86.46%, 69.50%, and 83.51%, respectively, and on NF-ToN-IoT from 3.82%, 6.17%, and 1.94% to 84.16%, 87.66%, and 83.37%. These results show that structured experience retrieval enables small, free, and locally deployable SLMs to achieve strong intrusion-detection performance, including performance exceeding that of the re- ported existing work used in our comparison.