arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.04756cs.CRcs.AI

PURPOSE:基于代理事实更新的RAG中毒冲突解决

PURPOSE: Poisoning Conflict Resolution in RAG via Proxy-Fact-Grounded Updates

Zijian Wang, Yubo Zhu, Muzhi Dong, Yanjun Lou, Yisheng Li, ZiLiang Zhang, Wei Tong, Yuan Zhang, Jingyu Hua, Sheng Zhong

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出PURPOSE黑盒中毒攻击,通过基于代理事实的非矛盾性更新实现对RAG冲突解决机制的有效攻击,在多数实验设置中攻击成功率显著优于现有方法。

中文摘要 AI 辅助

在检索增强生成(RAG)中,检索后冲突解决机制会对存在噪声或相互矛盾的检索段落进行仲裁。然而,该安全机制针对知识中毒的鲁棒性尚未得到充分研究。现有黑盒中毒方法均通过注入与冲突解决机制已确定内容正面矛盾的信息来实现目标答案,而这正是冲突解决机制原本要检测的信号。本文提出PURPOSE,一种严格的黑盒中毒攻击方法,将注入操作重新定义为最小化冲突的更新,而非提出反主张。PURPOSE会提取与查询相关的、近似冲突解决机制可能参考的事实,随后在这些事实中确立一个枢轴事件,以确保注入内容与冲突解决机制可验证的内容保持一致,同时引导生成器生成目标答案。在三个问答基准、五个生成器以及三种冲突解决方法的实验中,PURPOSE在45种设置中的35种达到了最高攻击成功率(ASR),且平均ASR较最强的现有攻击提升了9.7个百分点。这些结果表明,本文提出的中毒方法对RAG中的冲突解决机制是有效的,并确认非矛盾性注入是提升中毒攻击的可行模式。

英文摘要

In Retrieval-Augmented Generation (RAG), post-retrieval conflict resolution arbitrates among noisy or contradictory retrieved passages. However, the robustness of this safeguard against knowledge poisoning has not been adequately studied. Existing black-box poisoning methods all assert the target answer in frontal contradiction with what the resolver treats as settled, the very signal these methods are built to detect. We propose PURPOSE, a strict black-box poisoning attack that reframes the injection as an update that minimizes conflict, rather than as a counter-claim. PURPOSE extracts query-related facts approximating the resolver's possible reference, then grounds a pivot event in them to keep the injection consistent with what the resolver might verify while steering the generator toward the target answer. Across three QA benchmarks, five generators, and three conflict-resolution methods, PURPOSE attains the highest attack success rate (ASR) in 35 of 45 settings and exceeds the strongest prior attack with +9.7 mean ASR points. These results show that our poisoning method is effective against conflict resolution in RAG and identify non-contradicting injection as a practical mode to enhance poisoning attack.

发表机构

  • School of Computer Science, Nanjing University(南京大学计算机科学学院)
  • State Key Laboratory for Novel Software Technology, Nanjing University(南京大学计算机软件新技术国家重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

↑