arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.19430cs.CRcs.AIcs.MA

通道卫士:安全模型无法组合成安全的多智能体系统

ChannelGuard: Safe Models Do Not Compose into Safe Multi-Agent Systems

Elias Hossain, Md Mehedi Hasan Nipu, Fatema Tuj Johora Faria, Tasfia Nuzhat Ornee, Maleeha Sheikh

首次发表
浏览论文内容

中文总结 AI 辅助

研究多智能体大语言模型应用程序中智能体间通道安全问题,提出ChannelGuard深度防御框架,在通道设信息瓶颈门,通过文本与对抗短语库评分处理信息,能有效阻止工具中毒攻击,降低提示注入攻击成功率,保持准确率,白盒自适应释义可规避嵌入门。

中文摘要 AI 辅助

多智能体大语言模型应用程序由规划器、工作智能体、验证器和合成器组成,智能体之间的每一跳都是一个未受监控的通道,对手可借此偷运指令。现有防御措施仅保护输入边界或在应用程序外部运行。研究表明,在对八个攻击家族、五种防御措施和三种模型后端进行的2100次跟踪评估中,未受防御的管道在标准报告下看似完全安全,但其安全性几乎完全归功于云提供商的服务器端过滤器。结果报告掩盖了这种依赖性。提出ChannelGuard,一个无需训练的深度防御框架,在每个智能体间通道上设置信息瓶颈门,通过嵌入相似度对通道文本与对抗性短语库进行评分,确定地通过、压缩或阻止文本,不增加大语言模型调用。ChannelGuard的工具输出门在应用层阻止了30次工具中毒攻击,在不同后端表现相同,还将提示注入攻击成功率降低一半,同时保持GSM8K准确率不变。白盒自适应释义可规避每个嵌入门,而扰动投票基线表现更好。附录增加了基线、消融、扫描、良性保存分析和法官审计。

英文摘要

Multi-agent LLM applications chain a planner, worker agents, a verifier, and a synthesizer, and every hop between agents is an unmonitored channel through which an adversary can smuggle instructions. Existing defenses guard only the input boundary (IBProtector, Llama Guard, perplexity filters, SmoothLLM) or run outside the application as opaque, stochastic provider-side filters. We show this gap carries a consequence rarely measured: on a 2,100-trace evaluation across eight attack families, five defenses, and three model backends, an undefended pipeline that appears fully safe under standard reporting (attack success 0.000 on tool- and memory-poisoning) owes that safety almost entirely to the cloud provider's server-side filter (54 of 60 blocks on Azure GPT-5), and silently shifts to the agent model's own alignment on a backend without such a filter. Outcome-only reporting hides this dependence. We present ChannelGuard, a training-free defense-in-depth framework placing information-bottleneck gates on every inter-agent channel; each scores channel text against an adversarial phrase bank by embedding similarity and deterministically passes, compresses, or blocks it, adding no LLM call, while an attribution method records which layer stopped each attack. ChannelGuard's tool-output gate blocks Tool Poisoning 30 of 30 at the application layer, identically across Azure GPT-5, Anthropic Sonnet 4.5, and Anthropic Haiku 4.5, whereas the undefended pipeline shifts entirely across backends; it also lowers Prompt Injection attack success by half (0.333 to 0.167) and preserves GSM8K accuracy exactly (0.867). White-box adaptive paraphrase evades every embedding gate, where a perturb-and-vote baseline does better. An extended appendix adds baselines, ablations, sweeps, a benign-preservation analysis, and a judge audit (kappa = 0.900), at a total cost of 47.36 USD.

发表机构

  • College of Engineering and Computer Science, University of Central Florida(工程与计算机科学学院,中央佛罗里达大学)
  • Department of Computer Science and Engineering, North South University(计算机科学与工程系,北南大学)
  • Computer Science and Engineering, Ahsanullah University of Science and Technology(计算机科学与工程,阿沙努拉科学与技术大学)
  • Department of Electrical and Computer Engineering, Purdue University Fort Wayne(电气与计算机工程系,普渡大学弗拉特沃恩分校)

机构由 AI 辅助整理,请以论文原文为准。

↑