arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

覆盖并非包含:针对向量检索协同投毒的准入时刻防御的根本局限

Coverage Is Not Containment: A Fundamental Limit of Admission-Time Defenses Against Coordinated Poisoning of Vector Retrieval

Prashant Kumar Pathak, Tarun Kumar Sharma

arXiv 2608.16044首次发表:更新:

AI 中文总结

该研究揭示摄入时刻防御无法抵御向量检索协同投毒,提出检索时刻检测器可捕获全部攻击,证明覆盖并非包含,稳健防御需转向查询需求。

AI 中文摘要

检索增强生成(RAG)通过从向量存储中检索段落并将其作为上下文信任来回答问题,因此任何能够添加文档的人都可能试图引导答案。一种近期颇具吸引力的防御方法是在数据摄入时过滤投毒,拒绝任何表现为枢纽的文档。我们证明,这种防御以及所有摄入时刻过滤器都会被协同攻击者击败,该攻击者注入少量单独来看无异常的文档,这些文档共同围绕一个目标查询并占据其 top-k(在 BGE-large / BEIR 上,m=10 个文档占据 10/10;在实时 HNSW 索引上为 9.9/10)。该攻击并非理论性的,其被实现为普通流畅文本并通过 BGE-large + HNSW + Qwen2.5-7B 管道端到端运行时,会使生成器在 88% 的目标中输出攻击者植入的声明,而未注入时这一比例为 0%。且没有任何准入时刻防御能阻止它:在数据摄入时,攻击锥在几何上与合法小众上传完全相同,因此——直接测量这一点——最强的训练分类器在获得所有特征和数千个示例的情况下,区分两者的能力不优于随机,在 1% 的假阳性率下仅能捕获 4.2% 的攻击。我们针对整个摄入时刻统计类别(仅基于文档和参考查询做出的任何决策)证明了这一局限,且该局限在两个语料库和五个编码器上均会重现并恶化。区分攻击与合法小众摄入的唯一信号是查询的需求,而该需求在检索前不可见,这也是攻击的逃脱途径:在相同 1% 的假阳性率下,观察需求的检索时刻检测器能捕获 100% 的攻击。准入门对查询空间的覆盖并非对协同投毒的包含;稳健防御必须越过前门,转向需求。

英文摘要

Retrieval-augmented generation (RAG) answers a question by retrieving passages from a vector store and trusting them as context, so anyone who can add documents can try to steer the answer. A recent, appealing defense filters poisoning at ingestion, rejecting any document that behaves like a hub. We show it -- and every ingestion-time filter -- is defeated by a coordinated adversary that injects a handful of individually unremarkable documents which together surround one target query and seize its top-k (on BGE-large / BEIR, m=10 documents take 10/10; 9.9/10 on a live HNSW index). The attack is not theoretical. Realized as ordinary fluent text and run end-to-end through a BGE-large + HNSW + Qwen2.5-7B pipeline, it makes the generator emit the attacker's planted claim in 88% of targets, versus 0% without the injection. And no admission-time defense stops it: at ingestion an attack cone is geometrically identical to a legitimate niche upload, so -- measuring this directly -- the strongest trained classifier, given every feature and thousands of examples, separates the two no better than chance, catching 4.2% of attacks at a 1% false-positive rate. We prove this limit for the entire class of ingestion-time statistics (any decision from documents and reference queries alone), and it reproduces -- and worsens -- across two corpora and five encoders. The one signal that separates an attack from legitimate niche ingestion -- a query's demand -- is invisible before retrieval, which is also the escape: a retrieval-time detector that observes demand catches 100% of the attacks at the same 1% false-positive rate. Coverage of the query space by an admission gate is not containment of coordinated poisoning; robust defense must move past the front door, to demand.

Comments10 pages, 9 figures. Preprint; under submission

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑