arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RAG-CT:通过扫描提示分布缓解检索增强生成系统中的隐私风险

RAG-CT: Mitigating Privacy Risks on Retrieval-Augmented Generation Systems via Scanning Prompt Distribution

Xingyu Lyu, Jiayimei Wang, Jianfeng He, Ning Wang, Yidan Hu, Yimin Chen

arXiv 2609.16095首次发表:更新:

发表机构

University of Massachusetts Lowell; City University of Hong Kong; Virginia Tech; University of South Florida; Rochester Institute of Technology(马萨诸塞大学洛厄尔分校; 香港城市大学; 弗吉尼亚理工大学; 南佛罗里达大学; 罗切斯特理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对RAG系统易受PII泄露攻击的问题,提出基于熵和边际分布分析的轻量级防御方法RAG-CT,在多个数据集上显著降低泄露并优于现有防御。

AI 中文摘要

检索增强生成(RAG)已成为一种强大的范式,通过将响应基于外部知识来提升大型语言模型(LLMs)生成内容的质量,从而减少幻觉和事实错误。然而,近期研究突出了一个关键漏洞:攻击者可以利用检索过程从底层语料库中提取个人身份信息(PII)。为缓解这一风险,我们提出了一种新颖的防御方法RAG-CT,该方法通过分析查询的熵和边际分布,并采用基于评分的检测技术来识别恶意查询。我们在两个数据集上,针对四种最先进的攻击策略和四种防御基线进行了大量实验,结果表明我们的方法显著减少了PII泄露,同时优于现有防御手段。这项工作提供了一种轻量级且有效的机制,无需修改底层LLM或检索器即可保护RAG系统免受PII泄露。

英文摘要

Retrieval-Augmented Generation (RAG) has emerged as a powerful paradigm for improving the quality of generated contents of Large Language Models (LLMs) by grounding responses in external knowledge, thus reducing hallucinations and factual errors. However, recent studies have highlighted a critical vulnerability: adversaries can exploit the retrieval process to extract personally identifiable information (PII) from the underlying corpus. To mitigate this risk, we propose a novel defense, RAG-CT, that identifies malicious queries by analyzing their entropy and margin distributions and using a score-based detection method. Extensive experiments with four state-of-the-art attack strategies and four defense baselines on two datasets show that our approach significantly reduces PII leakage while outperforming existing defenses. This work provides a lightweight yet effective mechanism to protect RAG systems against PII leakage without requiring modifications to the underlying LLM or retriever.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑