arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ToxicRAG:通过单次知识投毒攻击破坏检索增强生成系统

ToxicRAG: Compromising Retrieval-Augmented Generation Systems via Single-Shot Knowledge Poisoning Attacks

Haozhe Lu, Jiaqi Li, Xinyuan Zhu, Xiang Li

arXiv 2609.11082首次发表:更新:

发表机构

School of Software and Microelectronics, Peking University; College of Cryptology and Cyber Science, Nankai University(北京大学软件与微电子学院; 南开大学密码学与网络科学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ToxicRAG提出单文档知识投毒攻击,通过叙事式虚假信息更新使RAG系统输出错误答案,在多个数据集和模型组合上达到0.61-0.91的攻击成功率,超过现有基线。

AI 中文摘要

检索增强生成(RAG)可以将大语言模型(LLM)的输出基于外部证据,但这也使系统面临知识投毒的风险。具有代表性的攻击使用多个注入文档或模板,直接断言目标答案。我们提出了ToxicRAG,一种每目标单文档的攻击方法,将错误信息表达为连贯的知识更新叙事。生成的文档首先承认先前被接受的答案,引入看似使其无效的虚构事件,然后将攻击者选择的答案归因于一组所谓的权威机构。一个以答案为中心的自我验证循环在替代语言模型未复现目标答案时,可选地修订候选文档。我们在Natural Questions、HotpotQA和MS-MARCO各100个目标问题上评估了该攻击,使用了四个受害大语言模型和四个密集检索器。在本文报告的采样语料库设置中,ToxicRAG在十二个数据集-模型组合上的攻击成功率(ASR)介于0.61至0.91之间。它在每个组合中达到或超过最强评估基线,差距范围为0至11个百分点。这些结果表明,在评估的RAG配置下,叙事形式的投毒文档仍能保持影响力,并激励对RAG系统中事实一致性和来源溯源性的进一步研究。

英文摘要

Retrieval-Augmented Generation (RAG) can ground large language model (LLM) outputs in external evidence, but it also exposes the system to knowledge poisoning. Representative attacks use multiple injected documents or templates that directly assert a target answer. We present ToxicRAG, a one-document-per-target attack that expresses misinformation as a coherent knowledge-update narrative. The generated document first acknowledges the previously accepted answer, introduces fabricated events that appear to invalidate it, and then attributes the attacker-selected answer to a set of purported authorities. An answer-focused self-validation loop optionally revises a candidate when a surrogate language model does not reproduce the target answer. We evaluate the attack on 100 target questions from each of Natural Questions, HotpotQA, and MS-MARCO, using four victim LLMs and four dense retrievers. In the sampled-corpus setting reported in this paper, ToxicRAG obtains ASRs between 0.61 and 0.91 across the twelve dataset--model combinations. It matches or exceeds the strongest evaluated baseline in every combination, with margins ranging from 0 to 11 percentage points. These results show that narrative-form poisoned documents can remain influential under the evaluated RAG configurations and motivate further study of factual consistency and source provenance in RAG systems.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑