arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24994cs.ITcs.CRmath.IT

反馈编码实现推理时隐蔽智能体通信

Feedback Coding Enables Inference-Time Covert Agentic Communication

Sidong Guo, Sajani Vithana, Atefeh Gilani, Lalitha Sankar, Oliver Kosut, Flavio P. Calmon

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出BAM反馈编码方案,将黑盒LLM隐写术视为带反馈的序列通信问题,通过后验匹配与解码确认实现低错误率隐蔽通信,在8位载荷下错误率远低于现有基线。

中文摘要 AI 辅助

随着大型语言模型(LLMs)越来越多地被用于自动化数字交互,用户可以利用LLM生成的文本作为掩护,在看似无害的对话中进行隐蔽通信。然而,现有的LLM隐写术主要是白盒的,要求发送方和接收方共享覆盖统计信息,通常通过访问模型权重和提示来实现。黑盒方案通过允许接收方仅基于生成的文本进行操作来消除这一要求,但当前方法依赖于固定长度、开环的水印技术,在变长令牌生成下会遭受高解码错误率。我们将黑盒LLM隐写术重新表述为一个具有因果、无噪声反馈的序列通信问题:每个生成的令牌都被双方观察,并可指导后续嵌入。基于这一视角,我们引入了Burnashev自适应后验匹配(BAM),一种结合后验匹配与解码确认阶段的反馈编码方案。该设计灵感来自经典的信息论反馈编码原理,而其安全性通过密码学归约证明得以确立。在三个开放权重语言模型上,我们证明BAM在约50个令牌内、1000次试验中,对8位载荷实现了0-0.1%的经验消息错误率,而最强的黑盒基线在相似长度下错误率为10-17%。基于所提出的隐写算法,我们展示了端到端通信协议的可行性,该协议在多种对话设置中实现了高通信速率。

英文摘要

As large language models (LLMs) are increasingly used to automate digital interactions, users can leverage LLM-generated text as cover for covert communication within seemingly benign conversations. Existing LLM steganography, however, is predominantly white-box, requiring the sender and receiver to share the cover statistics, typically through access to the model weights and prompt. Black-box schemes remove this requirement by allowing the receiver to operate solely on the generated text, but current approaches rely on fixed-length, open-loop watermarking techniques that suffer from high decoding error rates under variable-length token generation. We recast black-box LLM steganography as a sequential communication problem with causal, noiseless feedback: every generated token is observed by both parties and can guide subsequent embedding. Based on this perspective, we introduce \textbf{B}urnashev \textbf{A}daptive Posterior \textbf{M}atching (BAM), a feedback-coding scheme that combines posterior matching with a decode-and-confirm phase. The design is inspired by classical information-theoretic feedback-coding principles, while its security is established through a cryptographic reduction proof. Across three open-weight language models, we demonstrate that BAM attains 0-0.1\% empirical message error on an 8-bit payload in around 50 tokens, across 1000 trials, versus 10-17\% for the strongest black-box baseline at comparable length. Building on the proposed steganography algorithm, we demonstrate the feasibility of an end-to-end communication protocol that achieves high communication rates across multiple conversational settings.

发表机构

  • Georgia Institute of Technology(佐治亚理工学院)
  • Harvard University(哈佛大学)
  • Arizona State University(亚利桑那州立大学)

机构由 AI 辅助整理,请以论文原文为准。

↑