arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RAGSentinel:用于鲁棒检索增强生成的可验证几何共识

RAGSentinel: Certifiable Geometric Consensus for Robust Retrieval-Augmented Generation

Yueyang Quan, Anjun Gao, Yufei Xia, Minghong Fang, Zhuqing Liu

arXiv 2608.23965首次发表:更新:

发表机构

University of North Texas; University of Louisville(北得克萨斯大学; 路易斯维尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对RAG系统的中毒攻击漏洞,提出无训练无标签的RAGSentinel防御方法,通过几何共识过滤中毒文档,实验显示其能低攻高保准且抗自适应攻击。

AI 中文摘要

检索增强生成(RAG)通过将模型响应基于外部文档进行 grounding,提升了大语言模型的事实性,但同时也暴露了一个关键安全漏洞:注入到知识库中的对抗性文档可进入上下文窗口,引导模型生成目标错误答案。现有的检索后防御措施依赖指令遵循、参数知识或文本级一致性,均可能被自适应攻击者模仿或优化规避。我们提出RAGSentinel,一种针对黑盒RAG系统的无训练、无标签防御方法。RAGSentinel利用代理编码器测量检索文档引发的查询条件隐态偏移,去除共享主题方向,并从鲁棒多数共识中过滤出作为几何异常值的中毒文档。我们证明,在诚实多数假设和表示级分离条件下,RAGSentinel可精确恢复无中毒的多数规模上下文。在三个问答数据集、三个大语言模型家族及多种中毒攻击下的实验表明,RAGSentinel始终保持低攻击成功率,同时维持竞争力的准确率,且对拥有全流程知识的自适应攻击仍有效。

英文摘要

Retrieval-augmented generation (RAG) improves the factuality of large language models by grounding responses in external documents, but it also exposes a critical security vulnerability: adversarial documents injected into the knowledge database can enter the context window and steer the model toward targeted incorrect answers. Existing post-retrieval defenses rely on instruction following, parametric knowledge, or text-level consistency, all of which can be imitated or optimized against by adaptive attackers. We propose RAGSentinel, a training-free, label-free defense for black-box RAG systems. RAGSentinel uses a surrogate encoder to measure query-conditioned hidden-state shifts induced by retrieved documents, removes shared topic directions, and filters poisoned documents as geometric outliers from a robust majority consensus. We prove that, under an honest-majority assumption and a representation-level separation condition, RAGSentinel exactly recovers a poison-free majority-sized context. Experiments across three question-answering datasets, three LLM families, and multiple poisoning attacks show that RAGSentinel consistently achieves low attack success rates while preserving competitive accuracy and remaining effective against adaptive attacks with full pipeline knowledge.

CommentsTo appear in EMNLP 2026 (Main Conference)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑