arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2606.09005cs.CRcs.CL

文档作者控制信号冒充:对RAG安全边界的低成本间接提示攻击

Document-Authored Control-Signal Impersonation: A Low-Cost Indirect Prompt Attack on RAG Safety Boundaries

  • Chengdu University of Information Technology(成都信息工程大学)

机构由 AI 辅助整理,请以论文原文为准。

Jianguo Zhu

更新

AI总结:

研究检索增强生成系统中文档文本冒充控制信号的安全漏洞,提出非命令式间接注入攻击方法DACSI,并在多个模型上验证其有效性。

AI中文摘要:

检索增强生成(RAG)系统通常将用户查询、检索文档、元数据、系统标签和任务指令序列化为一个自然语言提示。我们研究了这种设计中的源权威边界失效:攻击者撰写的检索文本可以冒充元数据、来源、权威或披露策略信号,这些信号对模型而言似乎是控制相关的。我们将这种模式称为文档作者控制信号冒充(DACSI)。DACSI是间接提示注入中一种非命令式、类似元数据的载荷子类。其核心教训很简单:文档作者标签是数据,而非策略。命令式注入要求模型忽略、覆盖或违反策略;而DACSI则询问当RAG提示渲染将可信和不可信文本合并到同一自然语言通道时,不可信的文档文本是否可能被错误归因于授权控制信号。我们在六种模型设置、提示压力水平、注入基线、信号分类、RAG中介管道、系统控制探针、源权威归因探针和合成金丝雀格式上评估了DACSI。我们按模型机制解释证据,而非将其视为六次同等重复:DeepSeek V4 Pro和Qwen3.5-397B提供了最清晰的正向提升,DeepSeek V4 Flash是高易感性设置,GPT-5.5和Gemini 3.1 Pro Low是具有选择性残留风险的强边界探针,而GLM-4.7是饱和泄漏边界案例。在这些机制中,DACSI值得单独评估,因为它使用无命令的元数据/来源/策略表面,遵循RAG特定的源权威路径,并对源/通道分离做出响应。源权威探针是行为归因证据,而非内部机制的证明。

英文摘要:

Retrieval-augmented generation (RAG) systems often serialize user queries, retrieved documents, metadata, system labels, and task instructions into one natural-language prompt. We study a source-authority boundary failure in this design: attacker-authored retrieved text can impersonate metadata, provenance, authority, or disclosure-policy signals that appear control-relevant to the model. We call this pattern Document-Authored Control-Signal Impersonation (DACSI). DACSI is a non-imperative, metadata-like payload subclass within indirect prompt injection. Its central lesson is simple: document-authored labels are data, not policy. Command-style injection asks the model to ignore, override, or violate policy; DACSI asks whether untrusted document text can be misattributed as an authorized control signal when RAG prompt rendering collapses trusted and untrusted text into the same natural-language channel. We evaluate DACSI across six model settings, prompt-pressure levels, injection baselines, signal taxonomies, RAG-mediated pipelines, system-control probes, a source-authority attribution probe, and synthetic canary formats. We interpret the evidence by model regime rather than as six equal replications: DeepSeek V4 Pro and Qwen3.5-397B provide the cleanest positive lift, DeepSeek V4 Flash is a high-susceptibility setting, GPT-5.5 and Gemini 3.1 Pro Low are strong-boundary probes with selected residual risks, and GLM-4.7 is a saturated leakage boundary case. Across these regimes, DACSI warrants separate evaluation because it uses a command-free metadata/provenance/policy surface, follows a RAG-specific source-authority path, and responds to source/channel separation. The source-authority probe is behavioral attribution evidence, not proof of an internal mechanism.

补充信息

↑