arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TriShieldRAG:一种针对检索增强生成中知识腐败的三环深度防御框架

TriShieldRAG: 3 Rings, One Blind Spot in Layered Defenses for Retrieval-Augmented Generation

Susil Kumar Mohanty, Rohit Patel, Kosuru Yuvaraj, Jeenal Chaudhary, Disha Singhania

arXiv 2607.23838首次发表:更新:

发表机构

Indian Institute of Technology Jodhpur(印度理工学院焦特布尔分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对检索增强生成中知识腐败问题,构建TriShieldRAG框架,设置摄取防护、检索评分器和跨语言模型共识阶段三个环,推导环2和环3起作用的条件,评估表明该框架能大幅降低攻击成功率并保持良性查询准确性。

AI 中文摘要

检索增强生成(RAG)使大语言模型能在查询时使用从外部知识库检索的文档回答问题,这使其适用于私人数据等,但也意味着模型答案的可信度取决于检索器提供的内容。若知识库接受多方写入,攻击者只需少量对抗性文档就能使模型给出错误答案。PoisonedRAG证明了这一点。我们构建了TriShieldRAG来缩小差距,在整个流程中设置了三个独立的环:摄取防护,通过词汇和统计中毒特征筛选文档;检索评分器,根据来源和一致性加权信任分数对检索集重新排序;跨语言模型共识阶段,对三个架构不同的语言模型进行投票,并在出现分歧时允许一次有界重新检索。我们推导了环2和环3预期起作用的条件。在针对原始PoisonedRAG中的非自适应攻击者、一个有5000篇文档的维基百科知识库以及10个目标问题进行评估时,完整流程将攻击成功率从约91%降至约13%,同时保持了良性查询的准确性。

英文摘要

Retrieval-Augmented Generation (RAG) grounds LLM answers in query-time retrieved documents, so reliability depends on what the retriever returns. PoisonedRAG (Zou et al., USENIX Security'25) showed five crafted documents mislead an undefended system in nearly 90% of cases, and that single-stage defenses give limited robustness. We propose TriShieldRAG, a three-layered framework: an Ingest Guard for document-level screening, a Retrieval Scorer for trust-aware re-ranking, and a Cross-LLM Consensus over three diverse models. We reasoned that collectively screening, re-ranking and validating retrieved evidence would give complementary protection, limiting the ability of poisoned documents to succeed through any single failure. We evaluate against non-adaptive and adaptive poisoning. Non-adaptively, on the full 2.68M-passage Natural Questions (NQ) corpus with the original PoisonedRAG attack, it cuts attack success from 79 +/- 1.0% to 1 +/- 0.0%. Adaptive attacks expose fundamental limits of layering. By changing only the document formatting, without modifying the poison text or accessing the retriever, the attacker reduces the Ingest Guard score from 0.500 to 0.000 and bypasses it on all 500 tested documents across three corpora. The remaining layers then give no protection: 62 +/- 0.8% attack success against a 56 +/- 2.5% undefended baseline on NQ, and 85 +/- 0.6% against 86 +/- 0.6% on HotpotQA. Layered defenses relying on the same retrieved evidence fail together: poisoned context misleads both re-ranking and consensus validation. Minority-poison thresholds prove corpus-dependent, at 0.214, 0.251 and 0.558 rather than the derived 0.5; a closed form we proposed for these failed a pre-registered prediction and is retracted. Cross-model agreement is misleading, reaching 0.96 while attack success approaches 99%. We release the framework, the evasion-certification methodology and artifacts.

Commentsv2: Adds an adaptive-attacker evaluation in which Ring 1 is fully evaded (500/500 documents, three corpora); scales to the full NQ, HotpotQA and MS-MARCO corpora; corrects Proposition 1, whose boundary is corpus-dependent (0.214/0.251/0.558) not 0.5; retracts a proposed closed form after a pre-registered prediction failed. The v1 headline 91%-to-13% result is withdrawn

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑