arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CamoDocs:针对检索增强语言模型的伪装文档中毒攻击

CamoDocs: A Poisoning Attack Against Retrieval-Augmented Language Models Using Camouflaged Documents

Jaewon Jung, Haizhong Zheng, Hongsun Jang, Jaeyong Song, Beidi Chen, Jinho Lee

arXiv 2608.28389首次发表:更新:

AI 中文总结

本研究提出CamoDocs伪装文档中毒攻击,将对抗性文档伪装在良性内容中,在多种RAG防御、LLM及专有模型上实现高攻击成功率,同时避免易被检测的查询重叠痕迹。

AI 中文摘要

检索增强生成(RAG)利用外部文档增强大型语言模型(LLM)的能力,但公共或用户可编辑的来源使RAG系统面临数据中毒风险:攻击者可注入恶意文档,引导模型输出特定目标答案。现有中毒攻击常依赖查询包含,将目标查询插入中毒文档以提升检索效果,但这会产生词汇和嵌入空间的人工痕迹,易被过滤。本文提出CamoDocs,一种避免直接包含查询的中毒攻击,通过将对抗性文档伪装在良性内容中实现。CamoDocs将合成的良性草稿与对抗性草稿分块,用分散令牌替换良性块中选定的令牌,以扩散中毒文档的嵌入,并应用连贯性过滤限制可读性下降。在7种RAG防御、3种开放权重LLM和3个基准测试中,CamoDocs实现了强劲的平均攻击成功率(ASR),同时避免了简单查询检测所利用的查询重叠人工痕迹。它对专有模型仍有效,在GPT-5.4-mini上实现平均ASR为61.80%,在Claude-Haiku-4.5上为55.09%。最后,研究表明,像TrustRAG这类重擦除的聚类防御可降低ASR,但会在NeoQA等依赖检索的基准测试中造成严重的效用下降。代码可在该URL获取。

英文摘要

Retrieval-augmented generation (RAG) augments LLMs with external documents, but public or user-editable sources expose RAG systems to data poisoning: attackers can inject malicious documents to steer outputs toward targeted answers. Existing poisoning attacks often rely on query inclusion, inserting the target query into poisoned documents to improve retrieval; however, this creates lexical and embedding-space artifacts that make them easy to filter. We propose CamoDocs, a poisoning attack that avoids direct query inclusion by camouflaging adversarial documents among benign content. CamoDocs chunks synthesized benign and adversarial drafts, replaces selected tokens in benign chunks with dispersion tokens that spread poisoned-document embeddings, and applies coherence filtering to limit readability degradation. Across seven RAG defenses, three open-weight LLMs, and three benchmarks, CamoDocs achieves strong average ASR while avoiding query-overlap artifacts exploited by simple query detection. It also remains effective against proprietary models, achieving average ASRs of 61.80% on GPT-5.4-mini and 55.09% on Claude-Haiku-4.5. Finally, we show that erasure-heavy clustering defenses such as TrustRAG can reduce ASR, but only with substantial utility drops on retrieval-dependent benchmarks such as NeoQA. Code is available at https://github.com/jaewonalive/CamoDocs.

CommentsAccepted to EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑