arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.05640cs.MA

CaMeLs 能对话吗?保护多智能体系统免受间接提示注入攻击

Can CaMeLs Talk? Securing Multi-Agent Systems Against Indirect Prompt Injection Attacks

  • University of Oxford(牛津大学)
  • MATS Research(MATS 研究院)
  • AI Sequrity Company(AI Sequrity 公司)

机构由 AI 辅助整理,请以论文原文为准。

James Peters-Gill, Avi Semler, Henning Bartsch, Ilia Shumailov, Christian Schroeder de Witt

AI总结:

本研究针对层级多智能体系统中CaMeL安全保证不组合的问题,提出multi-CaMeL协议,通过分离指令与数据通道保留来源,将攻击成功率降至0.0%,并平衡了实用性成本。

AI中文摘要:

间接提示注入攻击——嵌入在大语言模型处理内容中的恶意指令——仍然是安全部署工具使用型智能体的主要障碍。CaMeL [Debenedetti et al., 2025] 通过将受信任的控制流与不受信任的数据分离,并在运行时强制执行基于能力的安全策略,从而缓解了单个智能体面临的这一威胁。在本研究中,我们探讨了 CaMeL 的安全保证是否能在层级多智能体系统中组合,其中智能体将其他智能体作为工具调用。我们发现 CaMeL 的保证无法组合。我们构造了一个具体的提示注入攻击,尽管所有组成智能体都单独运行 CaMeL,该攻击仍然成功。我们的攻击利用了不受信任的数据可以被下游智能体重解释为受信任输入这一事实。随后,我们引入了 multi-CaMeL,一种智能体间通信协议,通过将受信任的自然语言指令与通过独立数据通道传递的不受信任数据分离,从而在智能体边界上保留数据来源。我们在 AssetOpsBench 上评估了 multi-CaMeL 的实用性,并在 MultiAgentDojo(我们通过将 AgentDojo 扩展到多智能体环境而开发的基准)上评估了其安全性与实用性的权衡。我们发现,与单个智能体 CaMeL 的 0.2% 和无 CaMeL 时的 12.9% 相比,multi-CaMeL 将攻击成功率(ASR)降低至 0.0%。multi-CaMeL 会产生实用性成本,但该成本随着模型能力的增强而呈下降趋势,并且对于最强的模型而言成本适中,这表明能力更强的模型能更好地适应协议所施加的约束。

英文摘要:

Indirect prompt injection attacks - malicious instructions embedded in content processed by large language models - remain a major obstacle to safely deploying tool-using agents. CaMeL [Debenedetti et al., 2025] mitigates this threat for an individual agent by separating trusted control flow from untrusted data and enforcing capability-based security policies at runtime. In this work, we investigate whether CaMeL's security guarantees compose in hierarchical multi-agent systems, where agents invoke other agents as tools. We find that CaMeL's guarantees do not compose. We construct a concrete prompt-injection attack that succeeds despite all constituent agents individually operating CaMeL. Our attack exploits the fact that untrusted data can be reinterpreted as trusted input by a downstream agent. We then introduce multi-CaMeL, an agent-to-agent communication protocol that preserves provenance across agent boundaries by separating trusted natural-language instructions from untrusted data passed through a distinct data channel. We evaluate multi-CaMeL's utility on AssetOpsBench and its security-utility tradeoff on MultiAgentDojo, a benchmark we develop by extending AgentDojo to the multi-agent setting. We find that multi-CaMeL reduces attack success rate (ASR) to 0.0%, compared with 0.2% for individual-agent CaMeL and 12.9% with no CaMeL. Multi-CaMeL incurs a utility cost, but this cost trends downward as model capability increases and is modest for the strongest models, suggesting that more capable models better accommodate the constraints imposed by the protocol.

补充信息

↑