arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TRIS:一种针对知识投毒的三层检索完整性筛选器

TRIS: A Tri-Layer Retrieval Integrity Sieve Against Knowledge Poisoning

Muhaimin Bin Munir, Akib Jawad Ononto, Nazia Shehnaz Joynab, Bhavani Thuraisingham, Latifur Khan

arXiv 2609.00470首次发表:更新:

发表机构

University of Texas at Dallas(德克萨斯大学达拉斯分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对RAG的知识投毒攻击,提出三层检索完整性筛选器TRIS,通过跨嵌入空间聚类、结构过滤和LLM一致性验证,大幅降低攻击成功率并恢复干净准确率。

AI 中文摘要

检索增强生成(RAG)将大型语言模型建立在外部语料库的基础上,但对检索到的文档的隐性信任创造了一个关键攻击面:PoisonedRAG表明,少量精心设计的段落可以主导密集检索,并将生成引导至攻击者选择的答案。我们提出了三层筛选器(Tri-Layer Sieve),这是一种中间件防御机制,通过三个步骤净化检索到的证据:一是使用独立的评判模型进行跨嵌入空间聚类;二是对触发器-有效载荷(Trigger-Payload)人工制品进行结构过滤;三是进行LLM一致性验证。该设计利用了检索阶段投毒的一个关键弱点:单个文档必须同时满足一个嵌入几何结构、一个内部触发器-有效载荷结构和一个生成目标——但很少能同时满足这三者,即使是针对通过释义规避该弱点的自适应攻击者,这种脆弱性仍然存在。在使用Contriever检索(k=50)的Natural Questions、HotpotQA和MS-MARCO数据集上,该筛选器将黑盒攻击成功率从67.0%/87.0%/64.0%降至3.0%/14.0%/4.0%;在启用第三层的情况下,将NQ上的白盒HotFlip攻击成功率从约74%降至27.8%,并使中毒文档的MRR降至0.000,同时在攻击下将干净准确率从13%-33%恢复至58%-76%。在能够感知架构并通过释义触发器规避结构过滤器的对手下,启用一致性层可将自适应攻击成功率减半(NQ上从32.0%降至15.0%),同时将干净准确率提高18个百分点,在实时检索下增加的延迟约为16-19秒/查询。

英文摘要

Retrieval-Augmented Generation (RAG) grounds large language models in external corpora, but implicit trust in retrieved documents creates a critical attack surface: PoisonedRAG shows that a handful of crafted passages can dominate dense retrieval and steer generation toward attacker-chosen answers. We present the Tri-Layer Sieve, a middleware defense that sanitizes retrieved evidence through cross-embedding-space clustering with an independent judge model, structural filtering of trigger-payload artifacts, and LLM consistency verification. The design exploits a key weakness of retrieval-stage poisoning: a single document must satisfy one embedding geometry, one internal Trigger-Payload structure, and one generation objective - rarely all three simultaneously, a fragility that persists even against an adaptive attacker who paraphrases around it. On Natural Questions, HotpotQA, and MS-MARCO with Contriever retrieval (k=50), the Sieve reduces black-box Attack Success Rate from 67.0/87.0/64.0% to 3.0/14.0/4.0%, mitigates white-box HotFlip attacks from ~74% to 27.8% on NQ with Layer 3 enabled, and drives poisoned-document MRR to 0.000, while restoring clean accuracy from 13-33% under attack to 58-76%. Under an architecture-aware adversary who paraphrases triggers to evade the structural filter, enabling the consistency layer halves adaptive ASR (32.0% to 15.0% on NQ) while raising clean accuracy by 18 points, at an added latency of ~16-19 s/query under live retrieval.

Comments15 pages, 2 figures, 10 tables. Accepted to Findings of EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑