arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

可检索但不可靠:检索增强生成中的攻击与防御研究综述

Retrieved But Not Reliable: A Survey on Attacks, and Defenses in Retrieval-Augmented Generation

Minh Tran, Cuong Dang, Tuc Nguyen, Khanh-Tung Tran, Minh Huynh Nguyen, Trinh Chau, Kien Le, Do Xuan Long, Jiahao Zhang, Fali Wang, Hoang D. Nguyen, Thanh Le, Suhang Wang

arXiv 2608.24977首次发表:更新:

发表机构

Faculty of Information Technology, University of Science; Vietnam National University, Ho Chi Minh City; Virginia Tech; Indiana University; University College Cork; National University of Singapore; VNU University of Engineering and Technology; The Pennsylvania State University(胡志明市科学大学信息技术学院; 越南国家大学胡志明市分校; 弗吉尼亚理工大学; 印第安纳大学; 科克大学学院; 新加坡国立大学; 越南国家大学工程技术大学; 宾夕法尼亚州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本综述针对检索增强生成(RAG)的安全与鲁棒性问题,形式化其各阶段威胁模型,分类攻击目标并梳理各阶段防御及评估方法,为该领域提供统一研究框架。

AI 中文摘要

检索增强生成(Retrieval-Augmented Generation, RAG)通过将大语言模型的输出锚定到外部知识来增强其性能,提升了输出的事实性并减少了幻觉。与此同时,检索增强的流程引入了新的鲁棒性与安全风险,包括语料库投毒、后门攻击、隐私泄露及公平性违反。尽管该领域进展迅速,现有综述在处理攻击者目标、威胁模型及RAG全流程各阶段防御方面仍存在局限。本综述提供了统一且感知流程的RAG鲁棒性概述,我们对语料库、检索器及生成器的威胁模型进行了形式化,并将攻击按三大目标组织:准确性、隐私与公平性。我们进一步从感知流程的角度综述了防御措施,涵盖检索、重排序、生成及溯源阶段。此外,我们总结了用于更深入评估和解释RAG鲁棒性的鲁棒性基准与可解释性方法。

英文摘要

Retrieval-Augmented Generation (RAG) enhances large language models by grounding outputs in external knowledge, improving factuality and reducing hallucinations. At the same time, the retrieval-augmented pipeline introduces new robustness and security risks, including corpus poisoning, backdoor attacks, privacy leakage, and fairness violations. Despite rapid progress in this area, existing surveys remain limited in their treatment of attacker objectives, threat models, and stage-specific defenses across the full RAG pipeline. This survey presents a unified and pipeline-aware overview of RAG robustness. We formalize threat models over the corpus, retriever, and generator, and organize attacks into three main objectives: accuracy, privacy, and fairness. We further review defenses from a pipeline-aware perspective, covering the retrieval, rerank, generation, and traceback stages. In addition, we summarize robustness benchmarks and explainability methods for more deeply evaluating and explaining RAG robustness.

Comments24 pages, 6 figures. Accepted to Findings of the Association for Computational Linguistics: EMNLP 2026. Peer-reviewed through ACL Rolling Review (ARR)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑