arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RAGSieve:用于检索增强生成中知识投毒检测的自参考局部对比方法

RAGSieve: Self-Referenced Local Contrast for Knowledge-Poison Detection in Retrieval-Augmented Generation

Xinlong Xu, Yoshua Y. Li

arXiv 2608.13010首次发表:更新:

发表机构

Nanjing University of Information Science and Technology; Meituan(南京信息工程大学; 美团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出RAGSieve框架,含RSQ与RSG两种局部对比方法,在多QA数据集和投毒构造上实现优异检测性能,联合部署可显著降低RAG攻击成功率且保留部分未投毒检索性能。

AI 中文摘要

检索增强生成(RAG)将外部语料库作为推理证据,允许注入的文档支持攻击者选定的主张。现有检测器依赖可信参考、特定攻击伪影或对语料库拓扑敏感的全局阈值。本文提出RAGSieve,一种自参考检测框架,其参考由被检测系统构建。RAGSieve-Query(RSQ)执行查询局部对比,对同一检索的排名6-20位的候选文档与排名前5的候选文档打分,以检测答案锚点集中和载体过渡;RAGSieve-Graph(RSG)执行语料库局部对比,在查询到达前,将每个文档的语义相似但词汇不同的邻居与其局部基线对比,以检测协同密度。在三个问答(QA)数据集和六种投毒构造上,RSQ的AUROC达95.2%,在5%干净文档移除率下检测到82.2%的投毒,而GMTP的对应值为81.1%/52.5%;RSG的对应值为93.3%/79.8%,而CleanBase的对应值为79.4%/37.6%。联合部署将攻击成功率从67.4%降至14.0%,同时在未投毒检索上保留41.3%的F1值,证明其无需投毒标签或可信语料库,即可在语料库摄入和查询阶段均提供实用保护。源代码可在该https URL获取。

英文摘要

Retrieval-augmented generation uses an external corpus as inference-time evidence, allowing an attacker to promote a false answer by injecting a handful of documents. Detection must distinguish this manipulation from ordinary relevance without knowing which queries or documents are targeted. Existing detectors use text irregularity, candidate consensus, or corpus-level graph structure, whose reliability varies with the attack and local context. We present RAGSieve, which constructs a reference matched to each detection scope. At query time, RAGSieve-Query (RSQ) compares generation candidates with the lower-ranked tail of the same retrieval, exposing answer-token concentration and carrier-payload seams. At corpus time, RAGSieve-Graph (RSG) compares each document's strongest semantic relations with its own neighborhood floor to measure coordinated density. Neither requires poison labels, a trusted corpus, or training. Across three QA datasets, three dense retrievers, and six poisoning constructions, RSQ reaches 95.2% AUROC and detects 82.2% of poison at a 5% clean-removal budget, against 81.1% and 52.5% for the strongest query-time baseline; RSG reaches 93.3% and 79.8% against 79.4% and 37.6% for the strongest corpus-time baseline, with a 79.6% versus 1.4% detection rate on camouflaged injections. Joint deployment cuts attack success from 67.4% to 16.1% while retaining unpoisoned-retrieval F1 at 41.0%, compared with 42.1% without filtering. Source code is available at https://github.com/XrazyMee/RAGSieve.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑