arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16818cs.CR

InceptionRAG:针对检索增强生成的隐蔽投毒攻击

InceptionRAG: Stealthy Poisoning Attack Against Retrieval-Augmented Generation

Jiachang Zhang, Min Chen, Xiao Ren, Zhenyong Zhang, Yuanchao Shu, Yunjun Gao, Zhikun Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出InceptionRAG,一种通过将恶意载荷分解为休眠段落并利用多跳推理诱导LLM自推错误信息的隐蔽投毒攻击,在三个数据集和三个LLM上攻击成功率超80%,并揭示更强推理能力反而增加脆弱性,同时提出HODOR防御。

中文摘要 AI 辅助

检索增强生成(RAG)系统通过外部知识增强大语言模型(LLMs)的能力,但已被证明易受语料库投毒攻击。现有的针对RAG的投毒攻击主要集中于单点显式注入,即恶意载荷完全封装在单个文档中。因此,最近的防御机制已发展到能有效识别并削弱这些威胁。在本文中,我们首先验证了现有防御机制对于一类新威胁——间接逻辑归纳——是不充分的。基于这一观察,我们提出了InceptionRAG,一种颠覆标准攻击范式的隐蔽攻击机制。InceptionRAG不是注入显式恶意载荷,而是将其分解为一系列休眠段落。这些段落看似无害,单独检查时可绕过现有防御机制。然而,当它们被一起检索到时,会触发LLMs通过多跳推理自行推导出目标错误信息。为了进一步提高InceptionRAG在黑盒设置中的适用性,我们提出了零阶后缀优化(ZOSO)来自动生成权威后缀。在三个数据集和三个LLM上的广泛评估表明,即使在严格的对抗约束下,InceptionRAG也能实现超过80%的攻击成功率。特别地,InceptionRAG展现出卓越的规避能力,能有效绕过那些缓解传统单文档注入的既有防御。我们的发现揭示了一个令人担忧的悖论:LLMs更强的推理能力增加了它们对基于推理的投毒攻击的脆弱性。为了减轻潜在的滥用,我们提出了一种基于文档隔离的防御方法HODOR,它解耦了对抗性逻辑依赖。

英文摘要

Retrieval-augmented generation (RAG) systems enhance large language models (LLMs) with external knowledge but have been demonstrated to be vulnerable to corpus poisoning. Existing poisoning attacks against RAG largely focus on single-point explicit injection, where the malicious payload is fully encapsulated within a single document. Consequently, recent mitigation mechanisms have evolved to identify and diminish these threats effectively. In this paper, we first verify that existing mitigation mechanisms are insufficient for a new class of threats: indirect logic induction. Motivated by this observation, we introduce InceptionRAG, a stealthy attack mechanism that subverts the standard attack paradigm. Instead of injecting explicit malicious payloads, InceptionRAG fragments it into a chain of dormant passages. These passages appear harmless and can bypass existing mitigation mechanisms when examined separately. However, when retrieved together, they trigger LLMs to self-deduce target misinformation via multi-hop reasoning. To further improve the applicability of InceptionRAG in black-box settings, we propose zeroth-order suffix optimization (ZOSO) to automate the generation of authoritative suffixes. Extensive evaluations across three datasets and three LLMs demonstrate that InceptionRAG achieves an attack success rate exceeding 80% even under rigorous adversarial constraints. In particular, InceptionRAG shows superior evasion capabilities, effectively bypassing established defenses that mitigate traditional single-document injections. Our findings expose a concerning paradox: the stronger reasoning capabilities of LLMs increase their vulnerability to reasoning-based poisoning attacks. To mitigate potential misuse, we propose a document isolation-based defense, HODOR, which decouples adversarial logical dependencies.

发表机构

  • Zhejiang University(浙江大学)
  • Vrije Universiteit Amsterdam(阿姆斯特丹自由大学)
  • Guizhou University(贵州大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑