防御检索增强型入侵检测系统免受知识投毒与提示注入攻击
Defending Retrieval-Augmented Intrusion Detection Against Knowledge Poisoning and Prompt Injection
浏览论文内容
中文总结 AI 辅助
本文提出RAG-IDS三层多智能体入侵检测框架,通过软信任评分、LECC和提示净化的检索边界防御,可在知识投毒与提示注入攻击下恢复分类质量,在CIC-UNSW-NB15数据集上表现出良好的防御效果。
中文摘要 AI 辅助
检索增强生成(RAG)通过从向量知识库中检索语义相似的历史流量,使大语言模型能够对网络流量进行分类并生成可读的事件报告。然而,检索层引入了知识投毒和提示注入攻击的漏洞。本文提出RAG-IDS,这是一种三层多智能体入侵检测框架,其结合了软信任评分、标签嵌入一致性检查(LECC)和提示净化的检索边界防御机制,旨在在检索层攻击下恢复分类质量。在CIC-UNSW-NB15数据集上的实验表明,相对于干净未防御性能,在1%投毒时恢复率R=1.0,在30%投毒时R=0.57,且干净性能开销可忽略。在提示注入下,多文档检索将标签翻转成功率限制在0.6%-2.4%,而单文档检索则为35%-55%。消融实验显示,LECC是鲁棒性的主要贡献因素,而基于软信任的降级优于硬过滤。经防御的RAG流程为入侵检测提供了可解释、抗攻击的基础,非常适合与高吞吐量分类器混合部署。
英文摘要
Retrieval-Augmented Generation (RAG) enables large language models to classify network flows and generate human-readable incident reports by retrieving semantically similar historical traffic from a vector knowledge base. However, the retrieval layer introduces vulnerabilities to knowledge poisoning and prompt-injection attacks. We present RAG-IDS, a three-tier multi-agent intrusion detection framework with a retrieval-boundary defense combining soft trust scoring, label-embedding consistency checking (LECC), and prompt sanitization, designed to recover classification quality under retrieval-layer attack. Experiments on CIC-UNSW-NB15 show recovery relative to clean undefended performance ranging from R=1.0 at 1% poisoning to R=0.57 at 30%, with negligible clean-performance overhead. Under prompt injection, multi-document retrieval limits label-flip success to 0.6-2.4%, compared with 35-55% for single-document retrieval. Ablation results show that LECC is the primary contributor to robustness, while soft trust-based demotion outperforms hard filtering. The defended RAG pipeline offers an explainable, attack-resilient foundation for intrusion detection, well suited for hybrid deployment alongside high-throughput classifiers.
发表机构
- Laurentian University(劳伦森大学)
- Centennial College(百年学院)
- The University of Osaka(大阪大学)
机构由 AI 辅助整理,请以论文原文为准。