arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

针对多跳检索增强生成智能体的显著性诱导:威胁与防御

Salience Induction against Multi-Hop RAG Agents: Threat and Defense

Xingfu Zhou, Pengfei Wang, Yuan Zhou, Wei Xie, Xu Zhou

arXiv 2607.17535首次发表:更新:

发表机构

National University of Defense Technology(国防科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究多跳检索增强生成智能体面临的新攻击面——显著性通道,提出显著性诱导方法并定义操作符类,构建管道,引入基准。通过实验评估多种模型和架构,展示该方法有效性,强调智能体RAG需防御显著性 - 相关性解耦。

AI 中文摘要

智能检索增强生成(RAG)系统越来越多地检索外部证据并协调工具用于知识密集型应用。在多跳问答中,智能体跨文档链接事实。现有防御集中在内容中毒(注入错误事实)和提示注入(嵌入指令)。本文识别出第三个攻击面:显著性通道,即使所有检索到的声明为真且无指令时,事实位置、强调、框架和语义接近度也能重定向推理。将显著性诱导形式化为保持真值的编辑,重定向多跳属性绑定同时保持检索痕迹语义完整。定义六个显著性编辑操作符类并构建迭代提议者 - 验证者管道。引入SalientWiki - MH,一个带诱饵注释的多跳基准。对五个前沿模型家族和三种智能体架构的评估显示出广泛的通用性。在30%编辑预算下,显著性诱导攻击成功率达83.3%;最强基线防御后攻击成功率为75.7%。无目标重写仅通过降低中性任务成功率来减少攻击。轻量级输入侧防御显著性归一化在标准攻击下将攻击成功率降至15.3%,在自适应攻击下降至23.6%。结果表明仅靠真实性和指令过滤不足:强大的智能体RAG还需要防范显著性 - 相关性解耦。

英文摘要

Agentic retrieval-augmented generation (RAG) systems increasingly retrieve external evidence and orchestrate tools for knowledge-intensive applications. In Multi-Hop question answering, agents chain facts across documents. Existing defenses focus on content poisoning, which injects false facts, and prompt injection, which embeds directives. We identify a third attack surface: the salience channel, through which fact position, emphasis, framing, and semantic proximity can redirect reasoning even when all retrieved claims are true and no instructions are present. We formalize Salience Induction as truth-preserving edits that redirect Multi-Hop attribute binding while leaving the retrieval trace semantically intact. We define six Salience-Editing operator classes and build an iterative proposer-verifier pipeline under factual and stealth constraints. We also introduce SalientWiki-MH, a decoy-annotated Multi-Hop benchmark. Evaluations across five frontier model families (GPT, Claude, Gemini, DeepSeek, and Qwen) and three agent architectures (ReAct, Reflexion, and tool-calling) show broad generalization. Under a 30% edit budget, Salience Induction achieves an 83.3% attack success rate; the strongest evaluated baseline defense leaves 75.7% post-defense ASR. Untargeted rewriting further reduces attacks only by degrading neutral task success. Our lightweight input-side defense, Salience Normalization, reduces attack success to 15.3% under standard attacks and 23.6% under an adaptive attack. These results show that truthfulness and instruction filtering alone are insufficient: robust agentic RAG also requires defenses against salience-relevance decoupling.

Comments18 pages, 4 figures, 12 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑