AI 中文总结
提出KidnapRAG,一种针对代理式RAG系统的黑盒顺序投毒攻击,通过三种角色文档(诱饵、链环、恶意指令)劫持多步推理链,实验表明在多种框架和基准上优于现有方法。
AI 中文摘要
检索增强生成(RAG)系统容易受到投毒攻击,攻击者将恶意文档注入检索过程以操纵模型输出。最近的代理式RAG系统对此类攻击更具鲁棒性,因为它们迭代执行检索和推理,从而能够忽略弱相关的投毒文档并保持用户查询引发的推理链。然而,现有针对代理式RAG系统的攻击通常假设对系统提示、推理轨迹、检索器或模型参数具有白盒访问权限,这限制了它们在现实场景中的适用性。在本文中,我们研究针对代理式RAG系统的黑盒投毒攻击,其中攻击者只能发布外部可检索的投毒文档。我们提出KidnapRAG,一种顺序投毒攻击,利用三种角色文档(诱饵、链环和恶意指令)劫持代理的多步推理链,这些文档分别用于吸引初始检索、诱导查询重构和提供攻击者控制的证据。在多个代理式RAG框架、LLM骨干网络和基准上的实验表明,KidnapRAG在黑盒条件下始终优于现有的投毒基线。进一步分析表明,KidnapRAG逐步削弱原始检索意图,重定向检索行为,并增加对攻击者控制证据的依赖。我们的代码在此https URL公开。
英文摘要
Retrieval-Augmented Generation (RAG) systems are vulnerable to poisoning attacks that inject malicious documents into the retrieval process to manipulate model outputs. Recent Agentic RAG systems are more robust to such attacks because they iteratively perform retrieval and reasoning, allowing them to ignore weakly relevant poisoned documents and preserve the reasoning chain induced by the user query. However, existing attacks on Agentic RAG systems often assume white-box access to system prompts, reasoning traces, retrievers, or model parameters, limiting their applicability in realistic settings. In this paper, we study black-box poisoning attacks against Agentic RAG systems, where the attacker can only publish externally retrievable poisoned documents. We propose KidnapRAG, a sequential poisoning attack that hijacks the agent's multi-step reasoning chain using three role-specific documents: Bait, Chain-Link, and Mal-Ins, which attract initial retrieval, induce query reformulation, and provide attacker-controlled evidence, respectively. Experiments across multiple Agentic RAG frameworks, LLM backbones, and benchmarks show that KidnapRAG consistently outperforms existing poisoning baselines under black-box conditions. Further analyses show that KidnapRAG progressively weakens the original retrieval intent, redirects retrieval behavior, and increases reliance on attacker-controlled evidence. Our code is publicly available at https://github.com/chanwoochoi316/KidnapRAG.
CommentsAccepted to the Main Conference of EMNLP 2026