arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15660cs.CRcs.CL

CiteShade:多源检索增强生成中的引文洗白及其反事实防御

CiteShade: Citation Laundering in Multi-Source Retrieval-Augmented Generation and Its Counterfactual Defense

  • City University of Hong Kong(香港城市大学)

机构由 AI 辅助整理,请以论文原文为准。

Guo Fuzheng

AI总结:

本文提出首个针对RAG的引文洗白攻击CiteShade,通过优化构造恶意来源诱导模型错误归因,并验证反事实防御的有效性。

AI中文摘要:

检索增强生成(RAG)将语言模型的答案基于检索到的外部知识,并为每个答案附上标识其来源的引文。这些引文是用户的审计线索:它们使读者无需信任模型即可验证主张。先前关于RAG的安全研究关注攻击者能否篡改答案,而忽略了引文通道。我们证明该通道是一个新的且实用的攻击面。我们提出CiteShade,这是针对RAG的首个引文洗白攻击,其中控制单一来源的攻击者诱导模型生成攻击者选择的错误答案,并将其归因于不支持该答案的可信来源,同时正确答案的证据仍保留在上下文中。我们将该攻击形式化为优化问题,推导出三个必要条件(检索、生成和引文),并在无需任何指令的情况下构造满足这些条件的来源。在多源多跳问答中,该攻击将错误答案率从0.01提升至0.68,且来源删除证实恶意来源在每次测量中都是因果驱动因素。漏洞与模型引用倾向而非规模相关,在测试的最易引用模型上,显式指令下CLR达到0.84,无指令时达到0.64。我们进一步证明困惑度过滤和引文支持检查各自均不足,并提出一种反事实防御,用于验证实际驱动答案的来源。

英文摘要:

Retrieval-augmented generation (RAG) grounds a language model's answers on retrieved external knowledge and returns each answer with citations that identify its sources. Those citations are the user's audit trail: they let a reader verify a claim without trusting the model. Prior security work on RAG asks whether an attacker can corrupt the answer, leaving the citation channel unexplored. We show that this channel is a new and practical attack surface. We propose CiteShade, the first citation laundering attack to RAG, in which an attacker controlling a single source induces a model to produce an attacker-chosen wrong answer and to attribute it to a trusted source that does not support it, while the evidence for the correct answer remains in context. We formulate the attack as an optimization problem, derive three necessary conditions (retrieval, generation, and citation) and construct sources satisfying them without any instruction. On multi-source multi-hop question answering the attack raises the wrong-answer rate from 0.01 to 0.68, and source deletion confirms the malicious source is the causal driver in every measured case. Vulnerability tracks a model's propensity to cite rather than its scale, reaching CLR 0.84 under explicit instruction and 0.64 with no instruction at all on the most citation-prone model tested. We then show that perplexity filtering and citation-support checking are each insufficient, and propose a counterfactual defense that verifies which source actually drove the answer.

补充信息

↑