arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TopoGuard:基于图论的针对RAG分裂知识攻击的防御方法

TopoGuard: Graph Theory Based Defenses Against Split-Knowledge Attacks on RAG

Chahana Dahal, Zuobin Xiong

arXiv 2607.20437首次发表:更新:

发表机构

University of Nevada, Las Vegas(内华达大学拉斯维加斯分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对RAG系统分裂知识攻击的防御,提出基于图论的TopoGuard方法,通过构建语义相似性图检测恶意拓扑。实验表明其在捕获攻击、运行效率和对抗性方面优于现有方法,有效防御分裂知识攻击。

AI 中文摘要

生产检索增强生成(RAG)系统依靠聚合多个外部文档来回答复杂查询。然而,检索到的文档带来了新的威胁面,可能被用于发动分裂知识攻击。对手注入单独无害但组合并输入语言模型时会产生错误关联的文档。本文表明,这种新攻击对现有单文档过滤器(如LlamaGuard)在结构上不可见。为解决RAG中的此问题,引入TopoGuard,它通过从检索文档构建语义相似性图并检测具有恶意拓扑的上下文,专门针对分裂知识攻击。理论分析表明,即使输入有噪声,TopoGuard系列也有效且稳健。在两个检索数据集上进行了广泛实验,并与多种基线方法比较。例如,在HotpotQA数据集上,TopoGuard-$\lambda_2$+Entity在1%误报率下比LlamaGuard-2-8B多捕获21倍攻击(召回率32.6%对1.5%)。与使用大语言模型的生产RAG检测系统相比,TopoGuard变体在亚毫秒延迟下高效运行,在自适应对手和良性跨域查询下保持稳健。

英文摘要

Production Retrieval Augmented Generation (RAG) systems rely on aggregating multiple external documents to answer complex queries. However, the retrieved documents introduce a new threat surface that can be exploited to launch split-knowledge attacks. In this attack, the adversary injects documents that are individually benign but create false associations when combined and fed to language models. This paper shows that the new attack is structurally invisible to existing per-document filters, like LlamaGuard. To address this issue in RAG, this work introduces TopoGuard, a family of graph theory-based methods specifically targeting the split-knowledge attacks by building a semantic similarity graph from retrieved documents and detecting contexts with malicious topology. Grounded on the theoretical analysis, the TopoGuard family has been proven to be effective and robust even with noisy inputs. Extensive experiments are conducted on two retrieval datasets and compared with multiple baseline methods. Specifically, the TopoGuard-$λ_2$+Entity catches 21$\times$ more attacks than LlamaGuard-2-8B at 1\% FPR (32.6\% vs 1.5\% recall) on the HotpotQA dataset. Compared with production RAG detection systems using large language models, the proposed TopoGuard variants run efficiently at sub-millisecond latency and stay robust under adaptive adversaries and benign cross-domain queries.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑