arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多智能体系统中潜在通信的安全性

Safety of Latent Communication in Multi-Agent Systems

Muhammad Huzaifa, Sina Mavali, Thorsten Eisenhofer

arXiv 2609.39788首次发表:更新:

发表机构

CISPA Helmholtz Center for Information Security(CISPA赫尔姆霍兹信息安全中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究揭示潜在通信链接训练可增加有害遵从性,并提出强化学习攻击放大该效应,同时表明修复链接可降低风险,强调安全对齐需考虑系统整体。

AI 中文摘要

潜在通信使多智能体系统能够在内部表示空间中直接交换信息,从而减少基于文本通信的令牌、计算和延迟开销。为此,引入了轻量级可训练链接,将发送者的表示映射到接收者的输入空间。在本工作中,我们表明,即使良性的链接训练,相对于基于文本的通信,也可能增加有害遵从性,而底层经过安全对齐的智能体保持不变。攻击者可以通过在有害查询-响应对上优化链接,或投毒原本良性的训练数据来放大这种效应。我们进一步开发了一种强化学习攻击,该攻击在良性任务性能的同时奖励有害遵从性,而无需有害的目标响应。在三种通信拓扑和四个安全基准上,该攻击将平均有害遵从性得分从良性训练链接的27.9提高到76.9。与直接监督优化相比,它在两个良性效用基准上也实现了更高的平均准确率。将奖励调整为更安全的行为还可以修复受损的链接,在所有评估的攻击中显著降低有害遵从性,而无需更新智能体。总体而言,我们的结果表明,安全对齐需要考虑多智能体系统的整体。

英文摘要

Latent communication enables multi-agent systems to exchange information directly in internal representation space, reducing the token, computation, and latency overhead of text-based communication. To this end, lightweight trainable links are introduced to map the sender's representations into the receiver's input space. In this work, we show that even benign link training can increase harmful compliance relative to text-based communication while the underlying safety-aligned agents remain unchanged. An attacker can amplify this effect by optimizing the links on harmful query--response pairs or poisoning otherwise benign training data. We further develop a reinforcement-learning attack that rewards harmful compliance alongside benign task performance without requiring harmful target responses. Across three communication topologies and four safety benchmarks, this attack raises the mean harmful-compliance score from 27.9 with benignly trained links to 76.9. Compared with direct supervised optimization, it also achieves higher average accuracy on two benign utility benchmarks. Adapting the rewards toward safer behavior also enables repair of compromised links, substantially reducing harmful compliance across all evaluated attacks without updating the agents. Overall, our results show that safety alignment requires considering the multi-agent system as a whole. Code: https://github.com/Muhammad-Huzaifaa/latent-safety

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑