arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.28900cs.CRcs.MA

Codetta:高容量、无密钥且不可检测的多智能体共谋

Codetta: High-Capacity, Keyless, and Undetectable Multi-Agent Collusion

Qi Pang, Virginia Smith, Wenting Zheng

首次发表
浏览论文内容

中文总结 AI 辅助

Codetta提出一种高容量、无密钥且不可检测的隐写协议,使独立部署的LLM智能体在非对称设置中实现共谋,容量提升94倍,并指出审计需超越通信记录检查。

中文摘要 AI 辅助

基于大型语言模型(LLM)构建的多智能体系统正越来越多地部署在金融、医疗和软件工程等高风险场景中,这些系统中的智能体通过自然语言消息进行协调。然而,同样的通信渠道也让共谋的智能体能够泄露机密信息或协调未经授权的行动,而隐写术可以将此类通信隐藏在审计人员阅读记录时看似正常的输出中。现有的可证明不可检测的LLM隐写协议并不适用于实际部署。高容量方案假设一个对称设置,其中接收方能够复现发送方的输出分布;针对非对称智能体的最先进协议容量非常低;且大多数方法依赖预先共享的密钥。我们通过Codetta使不可检测的智能体共谋威胁具体化,这是一种针对现实非对称设置中独立部署智能体的高容量隐写协议。Codetta结合了共享的公共模型(用于估计通信信道)、保持发送方输出分布的采样机制以及自适应纠错码。它进一步通过隐写密钥交换消除了预先共享的密钥,使独立部署的智能体能够建立共享密钥,同时保持记录在计算上与普通模型输出不可区分。在三个智能体工作负载和三个发送方模型上,Codetta实现了最先进非对称协议容量的高达94倍,其密钥交换通过约8万可见令牌建立共享密钥,经验认证的失败概率至多为4.1×10⁻³。这些结果表明,有效不可检测的共谋在独立部署的智能体之间正变得可行,因此审计必须超越检查通信记录。

英文摘要

Multi-agent systems built on large language models (LLMs) are increasingly deployed in high-stakes settings such as finance, healthcare, and software engineering, where agents coordinate through natural-language messages. The same channels, however, let colluding agents exfiltrate confidential information or coordinate unauthorized actions, and steganography can hide such communication inside outputs that look ordinary to an auditor reading the transcript. Existing provably undetectable LLM steganography protocols are not suited to realistic deployments. High-capacity schemes assume a symmetric setting where the receiver can reproduce the sender's output distribution, the state-of-the-art protocol for asymmetric agents has very low capacity, and most approaches rely on a pre-shared secret key. We make the threat of undetectable agent collusion concrete with Codetta, a high-capacity steganographic protocol for independently deployed agents in realistic asymmetric settings. Codetta combines a shared public model that estimates the communication channel, a sampling mechanism that preserves the sender's output distribution, and an adaptive error-correcting code. It further removes the pre-shared key through a steganographic key exchange that lets independently deployed agents establish a shared key while keeping the transcript computationally indistinguishable from ordinary model outputs. Across three agent workloads and three sender models, Codetta achieves up to $94\times$ the capacity of the state-of-the-art asymmetric protocol, and its key exchange establishes a shared key with about 80k visible tokens at an empirically certified failure probability of at most $4.1\times 10^{-3}$. These results show that effectively undetectable collusion is becoming feasible between independently deployed agents, so auditing must go beyond inspecting communication transcripts.

发表机构

  • Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

↑