arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RAGuard:一种针对检索增强生成系统数据中毒的分层防御框架

RAGuard: A Layered Defense Framework for Retrieval-Augmented Generation Systems Against Data Poisoning

Pushkal Kumar, Tucker Nielson, Tanish Kolhe, Shubham Zala, Vincent Li

arXiv 2607.26339首次发表:更新:

AI 中文总结

该研究针对RAG系统的语料库中毒攻击,提出分层防御框架RAGuard,含对抗微调检索器与ZKIP两层,可将攻击成功率降至0,且保持检索性能,相关资源已开源。

AI 中文摘要

检索增强生成(Retrieval-Augmented Generation, RAG)系统将大型语言模型(Large Language Models, LLMs)与外部语料库结合,但这种依赖使其面临语料库中毒风险:恶意注入的段落会操纵检索证据。我们提出RAGuard,这是一种针对RAG流水线的事实性语料库中毒攻击的分层防御机制。第一层是在合成中毒文档(伪造事实、矛盾和推理陷阱)上对密集检索器进行对抗微调,使其在生成前对恶意段落进行降序排序。第二层是零知识推理补丁(Zero-Knowledge Inference Patch, ZKIP),它是一种无标签的黑盒过滤器:对于每个检索到的文档,它执行留一法解码,并通过移除该文档引发的语义偏移和输出熵变化对其进行评分。ZKIP不需要中毒标签、真实答案或模型内部结构访问权限,仅在反事实上下文中比较模型自身的答案。在中毒比例为5%至30%的中毒Natural Questions数据集上,仅对抗性检索器训练可降低但无法消除攻击成功率,而ZKIP在所有防御配置中将测得的攻击成功率降至0.000,同时保持Recall@5与干净语料库基线的差距在0.03以内。对Natural Questions和BEIR(NFCorpus)的监督分析证实,ZKIP所依赖的反事实信号携带可学习的中毒结构。该防御机制每个查询需要k+1次生成器传递(k=5时为6次);我们分析了可降低此开销的批处理和早停近似方法。我们还表明,保留关键词的中毒对BM25等词汇检索器基本无影响,这一观察明确了威胁模型的边界。代码、数据集和评估工具包已发布以支持可复现性。

英文摘要

Retrieval-Augmented Generation (RAG) systems ground large language models (LLMs) in external corpora, but this reliance exposes them to corpus poisoning: maliciously injected passages that manipulate retrieved evidence. We introduce RAGuard, a layered defense against \emph{factual} corpus-poisoning attacks on RAG pipelines. The first layer adversarially fine-tunes a dense retriever on synthetic poisoned documents (fabricated facts, contradictions, and reasoning traps), teaching it to downrank malicious passages before generation. The second layer, the Zero-Knowledge Inference Patch ZKIP, is a label-free, black-box filter: for each retrieved document, it performs a leave-one-out decode and scores the document by the semantic shift and output-entropy change that its removal induces. ZKIP requires no poison labels, no ground-truth answers, and no access to model internals; it compares the model's own answers under counterfactual contexts. On poisoned Natural Questions at 5--30\% poison ratios, adversarial retriever training alone reduces but does not eliminate attack success, while ZKIP drives the measured attack success rate to 0.000 in every defended configuration, keeping Recall@5 within 0.03 of the clean-corpus baseline. Supervised analyses on both Natural Questions and BEIR (NFCorpus) confirm that the counterfactual signals ZKIP relies on carry learnable poison structure. The defense costs $k{+}1$ generator passes per query ($6\times$ for $k{=}5$); we analyze batching and early-stopping approximations that reduce this overhead. We also show that keyword-preserving poisons leave lexical retrievers such as BM25 essentially unaffected, an observation that delineates the boundary of the threat model. Code, datasets, and evaluation harnesses are released for reproducibility.

CommentsAccepted to NeurIPS ResponsibleFM 2025, AAAI FrontierIR 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑