发表机构
Southern Illinois University(南伊利诺伊大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对RAG系统静态投毒攻击的盲区,提出自适应查询的RAG-NAROK框架,利用检索透明性生成源特定反驳文档,显著优于静态基线,揭示透明性与安全的矛盾。
AI 中文摘要
检索增强生成(RAG)系统已成为将大型语言模型(LLM)输出锚定于可验证外部知识的主流架构,但其对动态检索流程的结构性依赖引入了一类尚未充分探索的对抗性漏洞。现有知识库投毒攻击本质上是静态的:对抗性文档被预先计算并注入,而不考虑受害系统针对给定查询实际会检索到什么内容,这使得攻击无法感知生成器上下文窗口中围绕其载荷的竞争性文档环境。与对检索上下文视而不见的传统静态投毒攻击不同,我们提出RAG-NAROK(检索锚定生成否定与响应质量崩溃),一种自适应查询文本的RAG攻击框架。RAG-NAROK利用RAG流程固有的透明性,首先提取合法来源身份,然后生成锚定特定反驳文档,这些文档明确点名并贬低被检索来源,同时利用时效性和权威性偏差将文本生成引导至目标答案。我们的结果表明,RAG-NAROK在多个不同领域显著优于静态基线,揭示了RAG透明性与AI安全之间的根本性矛盾。
英文摘要
Retrieval augmented generation (RAG) systems have emerged as the dominant architecture for grounding large language model (LLM) outputs in verifiable external knowledge, yet their structural reliance on a dynamic retrieval pipeline introduces a largely unexplored class of adversarial vulnerability. Existing knowledge-base poisoning attacks are fundamentally static. Adversarial documents are pre-computed and injected without any awareness of what the victim system will actually retrieve for a given query, leaving the attack blind to the competitive documentary landscape that surrounds its payload in the generator's context window. Unlike traditional static poisoning attacks that are blind to the retrieved context, we introduce RAG-NAROK (Retrieval-Anchored Generation Negation And Response Quality Collapse), a RAG attack framework that adapts to the query text. RAG-NAROK exploits the transparency inherent in RAG pipeline to first extract the legitimate source identities, then generate Anchor-Specific Refutation documents that explicitly name and devalue retrieved sources while leveraging recency and authority biases to steer the text generation toward a target answer. Our results demonstrate that RAG-NAROK significantly outperforms static baselines across diverse domains, revealing a fundamental tension between RAG transparency and AI security.