迈向更安全的检索增强生成(RAG):仅具备系统2思考能力的智能体可访问不可信文档
When Detection Does Not Guarantee Resistance: Reasoning and Poisoned Context in RAG
浏览论文内容
中文总结 AI 辅助
本研究针对RAG系统的知识投毒漏洞,提出仅具备系统2推理能力的智能体可访问不可信文档的安全原则,通过实证证明该原则能提升模型对被破坏证据的鲁棒性且无需严格隔离。
中文摘要 AI 辅助
检索增强生成(RAG)显著提升了大语言模型(LLM)的性能,但这些系统仍易受知识投毒攻击,即检索文档中的错误信息会影响模型的最终输出。值得注意的是,LLM可能正确检测到某文档包含错误信息,却仍会受其影响。现有研究通过“隔离原则”解决此漏洞,该原则负责最终答案合成的模型无法直接访问原始证据,尽管有效,但这种严格隔离会引入大量计算开销。本研究提出一种改进的安全原则:仅具备审慎的系统2推理能力的智能体可访问不可信文档。为评估该原则,我们引入新指标量化错误信息检测与下游影响之间的差异,随后在这些指标上对最先进的推理语言模型与标准语言模型进行实证比较。结果表明,具备推理能力的模型对被破坏证据的鲁棒性显著更强,且无需隔离原则施加的严格隔离。这些发现为我们的改进原则提供了实证支持,并为安全RAG系统设计提供了更实用的基础。
英文摘要
Retrieval-Augmented Generation (RAG) exposes large language models to knowledge-poisoning attacks, where misinformation injected into retrieved documents can influence model outputs. Prior work has shown that models may detect contradictory evidence yet still allow it to influence their responses, revealing a gap between monitoring and control. We investigate whether deliberative reasoning changes this relationship. Because attack success and poison detection alone do not reveal whether detected poison continues to influence the final answer, we use two complementary measures: Cordon Rate, the probability that a model detects poison and nevertheless produces a poison-aligned answer, conditioned on a non-poison-aligned no-RAG response; and Leakage Rate, the probability that poisoned context influences the answer despite an explicit instruction to ignore retrieved documents. Across 200 SciFact questions, we compare reasoning-disabled and reasoning-enabled configurations of DeepSeek-V4-Flash and Qwen3.6-Plus. For DeepSeek-V4-Flash, reasoning reduces Cordon Rate from 0.205 to 0.075 and Leakage Rate from 0.235 to 0.140, despite increasing Attack Success Rate from 0.233 to 0.298 and decreasing Poison Detection Rate from 0.965 to 0.665. Qwen3.6-Plus shows the same qualitative pattern. These results demonstrate that poison detection and downstream influence capture distinct aspects of poisoning behavior.
发表机构
- University of Isfahan(伊斯法罕大学)
机构由 AI 辅助整理,请以论文原文为准。