arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

我们信任RAG吗?衡量检索增强生成在文档投毒下的鲁棒性

In RAG We Trust? Measuring Robustness of Retrieval-Augmented Generation Under Post-Retrieval Context Tampering

Iliano Fasolino

arXiv 2609.09243首次发表:更新:

发表机构

University of Milan(米兰大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过因子实验测量Llama 3.1 8B在RAG文档投毒下的鲁棒性,发现实体替换攻击影响最大,模型主要反应为弃权而非产生新虚假信息。

AI 中文摘要

检索增强生成(RAG)将语言模型基于检索到的文档,这减少了幻觉,但创造了新的攻击面:如果检索到的文本被篡改,模型可能会重复虚假信息。我们研究了一个小型量化模型Llama 3.1 8B在其检索上下文的某部分被投毒时性能下降的程度。测试了三种破坏策略:实体替换、数字替换和否定,每种策略分别应用于三个检索段落中的零个、一个、两个或三个,在基于FEVER构建的事实核查任务上进行了588次运行的因子扫描。准确率从干净上下文中的77.9%下降到所有三个段落被破坏时的43.5%。实体替换翻转了在干净上下文上正确答案的最大份额。基于数字的破坏在投毒段落占少数时保持平稳,一旦它们成为多数则跃升,我们使用查询级自助法区间重新检查了这一模式。模型很少发明新的虚假信息;其主要反应是弃权(不执行),而未经支持生成的词汇重叠代理在攻击下下降而非上升。该研究是一项小规模测量,使用粗略的自动标签;在解码得到控制和更强裁决到位之前,我们将策略对比视为提示性的。

英文摘要

Retrieval-augmented generation (RAG) grounds a language model in retrieved documents, which reduces hallucination but creates a new attack surface: if retrieved text is tampered with, the model may repeat the falsehood. We study how much a small quantized model, Llama 3.1 8B, degrades when a fraction of its retrieved context is poisoned. Three corruption strategies are tested, entity swap, number swap, and negation, each applied to zero, one, two, or three of the three retrieved passages, over a factorial sweep of 588 runs on a fact-checking task built from FEVER. Accuracy falls from 77.9% on clean context to 43.5% when all three passages are corrupted. Entity swap flips the largest share of answers that were correct on clean context. Number-based corruption stays flat while poisoned passages are a minority and jumps once they form a majority, a pattern we re-check with query-level bootstrap intervals. The model rarely invents new falsehoods; its dominant reaction is to abstain, and a lexical overlap proxy of unsupported generation falls under attack rather than rising. The study is a small-scale measurement with coarse automated labels; we treat the strategy contrasts as suggestive until decoding is controlled and stronger adjudication is in place.

Comments4 figures. Preprint also available on Zenodo: https://doi.org/10.5281/zenodo.22980485

DOI:10.5281/zenodo.22980485

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑