arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越单模型注入:多智能体系统中提示注入的威胁模型与防御架构

Beyond Single-Model Injection: A Threat Model and Defense Architecture for Prompt Injection in Multi-Agent Systems

Rudrendu Kumar Paul, Sourav Nandy

arXiv 2609.22949首次发表:更新:

发表机构

Boston University; University of Texas at Austin(波士顿大学; 德克萨斯大学奥斯汀分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多智能体系统提示注入威胁,提出含14种攻击向量的威胁模型及四种架构防御,将注入成功率从31.2%降至4.2%。

AI 中文摘要

现有的提示注入研究聚焦于单模型聊天机器人场景,即攻击者通过精心构造的输入操纵单个大语言模型(LLM)。多智能体系统通过单模型设置中不存在的三种机制放大了这一威胁:智能体间消息传递创建了外围防御不可见的注入通道,共享工具访问使得权限提升能够跨越智能体边界,信任传播允许被攻破的智能体影响上游编排器。我们构建了一个威胁模型,枚举了四大类共14种攻击向量:通过用户输入的直接注入(3种向量)、通过工具输出的间接注入(4种向量)、通过消息传递的智能体间注入(4种向量)以及通过编排器操纵的级联注入(3种向量)。在包含6个智能体的生产代表性系统上测试全部14种向量,我们发现即使存在系统提示级防护栏,67%的智能体仍易受至少一种范围违规攻击,且通过工具输出的间接注入在43%的尝试中成功。四种架构防御将整体注入成功率从31.2%降至4.2%:带来源追踪的消息签名(智能体间注入降低91%)、智能体边界处的输入/输出净化(间接注入降低78%)、按智能体角色进行权限限定的工具访问(权限提升被完全消除)以及智能体间通信模式的异常检测(84%的级联尝试被捕获)。

英文摘要

Existing prompt injection research focuses on single-model chatbot scenarios, where an attacker manipulates one LLM through crafted input. Multi-agent systems amplify this threat through three mechanisms absent from single-model settings: inter-agent message passing creates injection channels invisible to perimeter defenses, shared tool access enables privilege escalation across agent boundaries, and trust propagation allows a compromised agent to influence upstream orchestrators. We construct a threat model enumerating 14 attack vectors across four categories: direct injection via user input (3 vectors), indirect injection via tool outputs (4 vectors), inter-agent injection via message passing (4 vectors), and cascading injection through orchestrator manipulation (3 vectors). Testing all 14 vectors against a 6-agent production-representative system, we find that 67% of agents are vulnerable to at least one scope violation even with system-prompt-level guardrails, and indirect injection via tool outputs succeeds in 43% of attempts. Four architectural defenses reduce overall injection success from 31.2% to 4.2%: message signing with provenance tracking (inter-agent injection down 91%), input/output sanitization at agent boundaries (indirect injection down 78%), privilege-scoped tool access per agent role (privilege escalation eliminated entirely), and anomaly detection on inter-agent communication patterns (84% of cascading attempts caught).

CommentsAccepted at the AIWILD Workshop, ICML 2026. Camera-ready version

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑