arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DenialRAG:通过嵌入参数化弃权的单文档RAG污染攻击

DenialRAG: Single-Document RAG Poisoning via Embedded Parametric Denial

Abay Zhurekbay, Tao Liu, Fan Li

arXiv 2608.02678首次发表:更新:

发表机构

Lawrence Technological University(劳伦斯理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出DenialRAG单文档RAG污染攻击,通过嵌入参数化弃权引导LLM生成错误答案,经多数据集、多模型及防御评估,其在Mistral-7B上攻击效果最优,凸显RAG污染风险的复杂性。

AI 中文摘要

检索增强生成(RAG)系统易受语料库污染攻击:攻击者向检索语料库插入特制文档,可引导底层大语言模型(LLM)生成攻击者选定的错误答案。现有单文档攻击通常避免在污染段落中明确提及并反驳正确答案。本文研究了一种互补设计,提出DenialRAG——一种单文档污染攻击,该攻击明确提及正确答案、对其进行弃权(不执行),并给出攻击者控制的解释以支持错误答案。通过将正确答案与对应的污染答案置于同一检索段落中,DenialRAG将冲突直接嵌入生成器可见的上下文。我们针对三种开放域问答数据集、四个厂商的八个目标LLM以及五种推理时防御措施,将DenialRAG与四种已发表的单文档污染攻击进行评估。结果显示,攻击有效性高度依赖模型:DenialRAG在所有三个Mistral-7B数据集上实现最高攻击成功率(ASR),并在其他多个目标LLM上保持有效性,而其他攻击在部分模型场景中占优。防御结果显示ASR有显著降低,但保护效果不均,每种防御在部分场景中仍残留ASR。组件级和跨模型分析进一步表明,嵌入的弃权是测试组件中影响最大的部分,且不同污染机制在不同模型组中的有效性下降速率不同。综上,这些结果表明RAG污染风险无法用单一攻击家族或单一目标模型完全表征。

英文摘要

Retrieval-augmented generation (RAG) systems are vulnerable to corpus poisoning: an attacker who inserts a crafted document into the retrieval corpus can steer the underlying large language model (LLM) toward an attacker-chosen wrong answer. Prior single-document attacks typically avoid explicitly naming and refuting the correct answer inside the poisoned passage. In this paper, we examine a complementary design and propose \emph{DenialRAG}, a single-document poisoning attack that explicitly names the correct answer, denies it, and presents an attacker-controlled explanation for favoring the wrong answer. By placing both the correct answer and the corresponding poisoned answer inside the same retrieved passage, DenialRAG embeds the conflict directly into the context seen by the generator. We evaluate DenialRAG against four published single-document poisoning attacks across three open-domain question-answering datasets, eight target LLMs from four vendors, and five inference-time defenses. The results show that attack effectiveness is strongly model-dependent: DenialRAG achieves the highest attack success rate (ASR) on all three Mistral-7B datasets and remains effective on several other target LLMs, while other attacks dominate in some model regimes. Defense results show meaningful ASR reductions but non-uniform protection, with each defense leaving residual ASR in some settings. Component-level and cross-model analyses further identify the embedded denial as the most influential tested component and show that different poisoning mechanisms lose effectiveness at different rates across model groups. Together, these results show that RAG poisoning risk cannot be fully characterized by a single attack family or a single target model.

CommentsSubmit to ACSAC 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑