arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LENS:检索增强生成中集合级投毒的整体效果弱于各部分之和

LENS: The Sum Is Worse Than the Parts for Set-Level Poisoning in Retrieval-Augmented Generation

Kaisheng Fan, Yishu Gao, Xunzhu Tang, Tegawend'e F. Bissyand'e, Weizhe Zhang

arXiv 2609.35155首次发表:更新:

发表机构

School of Cyber Science and Technology, Harbin Institute of Technology; SnT, University of Luxembourg; Department of New Networks, Peng Cheng Laboratory(哈尔滨工业大学网络空间安全学院; 卢森堡大学科学与技术学院; 鹏城实验室新网络部)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对RAG的集合级组合投毒漏洞,提出黑盒多智能体框架LENS,通过嵌套双循环构建单份合理、全集触发攻击的投毒文档,在多项基准上显著提升攻击效果,为RAG防御提供压力测试工具。

AI 中文摘要

检索增强生成(RAG)会聚合多份外部文档的证据,但这种联合整合方式催生了一种尚未被充分研究的漏洞:单份文档中不存在的攻击效果可通过集合级组合显现。现有协同攻击并未在冻结的单轮RAG中明确要求所有真子集都不足以触发攻击。我们对集合级组合投毒进行了形式化定义:这类攻击中的文档单份看似合理,但组合起来会将RAG输出导向目标答案,而任意真子集都无法单独诱导出目标结果。为构建此类攻击,我们提出LENS——一种生成器黑盒多智能体框架,将攻击构建转化为带约束的证据组合问题。LENS将目标推理分解为查询条件化的解释透镜与互补事实,随后采用嵌套双循环工作流,使转向效果集中在完整集合中,同时抑制子集泄露。外环负责规划解释透镜与语义角色;内环负责合成文档并执行反例引导的修复。在四个基准和三个生成器上的测试中,返回的攻击包实现了0.852的全集攻击成功率(ASR)和0.784的检索后ASR@5,而其最强真子集的ASR仅为0.069。在相同冻结清单上评估的构建基线相比,LENS将全尝试端到端严格率@5(E2E-Strict@5)从0.244提升至0.363,相对增益达48.8%。盲法人工审核发现,68.3%的返回攻击包同时包含错误目标、明确的答案标准偏移,且在原始语义下不存在目标蕴涵关系。在四种已发表的防御方法测试中,LENS取得了最高的防御后全尝试ASR@5,平均超出最强基线0.141。综上,这些结果确立了证据组合是RAG的一个独特安全边界,并将LENS定位为针对文档集推理防御方法的压力测试工具。

英文摘要

Retrieval-augmented generation (RAG) aggregates evidence from multiple external documents, yet this joint integration creates an underexamined vulnerability: attack effects absent in individual documents can emerge through set-level composition. Existing coordinated attacks do not explicitly enforce that every proper subset remains insufficient in frozen single-round RAG. We formalize set-level compositional poisoning, where documents designed to remain individually plausible jointly redirect RAG outputs to a target answer, while proper subsets fail to induce the target on their own. To construct such attacks, we propose LENS, a generator-black-box multi-agent framework that casts construction as constrained evidence composition. LENS factorizes target inference into a query-conditioned interpretation lens and complementary facts, then uses a nested dual-loop workflow to concentrate steering in the full set while suppressing subset leakage. The outer loop plans the interpretation lens and semantic roles; the inner loop synthesizes documents and applies counterexample-guided repair. Across four benchmarks and three generators, returned packets achieve 0.852 full-set ASR and 0.784 post-retrieval ASR@5, while their strongest proper subsets reach only 0.069. Against construction baselines evaluated on the same frozen manifest, LENS improves all-attempt E2E-Strict@5 from 0.244 to 0.363, a 48.8% relative gain. A blinded human audit finds that 68.3% of returned packets combine an incorrect target, a definite answer-criterion shift, and no target entailment under the original semantics. Across four published defenses, LENS attains the highest defended all-attempt ASR@5, exceeding the strongest baseline by 0.141 on average. Together, these results establish evidence composition as a distinct RAG security boundary and position LENS as a stress test for defenses that reason over document sets.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑