当上下文产生危害:通过文档级注意力崩溃检测RAG中毒
When Context Bites: Detecting RAG Poisoning via Document-Level Attention Collapse
浏览论文内容
中文总结 AI 辅助
针对RAG中毒攻击现有检测方法失效的问题,提出基于文档级注意力崩溃的轻量级检测框架D-SCAN,可在攻击未改变最终答案时也实现有效检测。
中文摘要 AI 辅助
检索增强生成(RAG)对于增强大型语言模型至关重要,但RAG日益易受中毒攻击,即注入对抗性文档以操纵生成器输出。先前方法依赖困惑度和一致性检查等输出侧信号检测此类攻击,但分析显示,蓄意攻击常引发假置信度,中毒输出的困惑度甚至低于良性输出,导致基于不确定性的检测失效。为应对该挑战,研究人员探索生成器内部动态,识别出称为注意力崩溃的独特特征:与良性生成中分散的注意力不同,受攻击生成中注意力集中于中毒文档,熵值下降。基于此发现,提出轻量级检测框架D-SCAN(文档级信号崩溃分析),该框架监测注意力动态以识别受攻击生成。在多个攻击基准上的大量实验证明了该方法的有效性,且D-SCAN甚至能在攻击未改变最终答案时检测到攻击,代码可在指定URL获取。
英文摘要
Retrieval-augmented generation (RAG) is indispensable for enhancing large language models. However, RAGs are increasingly susceptible to poisoning attacks, in which adversarial documents are injected to manipulate generator outputs. Previous methods rely on output-side signals such as perplexity and consistency checks to detect such attacks. Nevertheless, our analysis reveals that deliberate attacks often induce false confidence, where poisoned outputs exhibit even lower perplexity than benign ones, rendering uncertainty-based detection ineffective. To address this challenge, we explore the internal dynamics of the generator and identify a distinctive signature termed \textit{Attention Collapse}. Unlike the dispersed attention in benign generations, attacked generations exhibit a decrease in entropy as attention concentrates on poisoned documents. Building on these findings, we propose \texttt{D-SCAN} (Document-level Signal Collapse Analysis), a lightweight detection framework that monitors attention dynamics to identify attacked generations. Extensive experiments on multiple attack benchmarks demonstrate the effectiveness of our method. Moreover, D-SCAN can detect attacks even when they fail to alter the final answer. Code is available at https://github.com/yingtaoren/D-Scan.git.