arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19161cs.AIcs.CR

超越文本记录:检测潜在多智能体通信中的隐蔽协同

Beyond the Transcript: Detecting Covert Co ordination in Latent Multi-Agent Communication

Ramneet Kaur, Pradyumna Chari, Ramesh Raskar, Jugad Singh, Sumit Kumar Jha, Anirban Roy

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对语言模型智能体通过不可见潜在状态进行隐蔽协同的问题,提出可验证潜在对齐(VLA)框架,设计三层监控器与可引导性框架,在多智能体拍卖基准上实现了高效的隐蔽协同检测与缓解。

中文摘要 AI 辅助

语言模型智能体可通过公共文本记录中不可见的连续隐藏状态进行通信,这为隐蔽的有害协同创造了机会。我们引入可验证潜在对齐(Verifiable Latent Alignments, VLA),这是一种用于监控和引导这些私有通信通道的激活感知框架。对于每个被监控的决策,VLA 使用共享事件标识符将私有潜在状态记录和通道状态与产生的公共动作关联起来,从而实现匹配的因果分析。我们的第一个贡献是一个仅使用中性数据的三层监控器,结合了表示异常检测、反事实动作分布影响以及稀疏自编码器解释支持。我们的第二个贡献是一个可引导性框架,涵盖黑盒行为指令和白盒匹配中性反事实。我们的第三个贡献是在受控多智能体拍卖基准上进行的评估,涵盖同构和异构模型对、多智能体可扩展性以及干预有效性。当将文本和潜在协同行合并为正样本时,该序列监控器对同构智能体的受试者工作特征曲线下平均面积(AUROC)为 0.993,对异构对为 0.854。在 Qwen3-0.6B 拍卖中,当有 25 至 100 个竞标者时,监控相对于所有可能的有向对仅需要很小的归一化负载,而完整白盒引导实现了 100% 的出价分布恢复,并将协同低价行为降低了 47.3 个百分点。由于完整白盒引导会重播匹配的中性反事实,其精确恢复是构造性的合理性检查。总体而言,这项受控研究表明,无需在攻击示例上训练主监控器即可监控评估的私有通道攻击,并且在可获得匹配反事实访问权限时可缓解这些攻击。

英文摘要

Language-model agents can communicate through continuous hidden states that are invisible in public transcripts, creating opportunities for covert harmful coordination. We introduce Verifiable Latent Alignments (VLA), an activation-aware framework for monitoring and steering these private communication channels. For every monitored decision, VLA links the private latent-state record and channel status to the resulting public action using a shared event identifier, enabling matched causal analysis. Our first contribution is a neutral-only three-layer monitor combining representation anomaly detection, counterfactual action-distribution influence, and sparse-autoencoder interpretation support. Our second contribution is a steerability framework spanning black-box behavioral instructions and white-box matched-neutral counterfactuals. Our third contribution is an evaluation on a controlled multi-agent auction benchmark covering homogeneous and heterogeneous model pairs, many-agent scalability, and intervention effectiveness. The sequential monitor achieves mean area under the receiver operating characteristic curve (AUROC) of 0.993 for homogeneous agents and 0.854 for heterogeneous pairs when text- and latent-collusion rows are pooled as positives. In Qwen3-0.6B auctions with 25-100 bidders, monitoring requires only a small normalized load relative to all possible directed pairs, while full white-box steering achieves 100% bid-distribution recovery and reduces collusive low-bid behavior by 47.3 percentage points. Because full white-box steering replays the matched neutral counterfactual, its exact recovery is a sanity check by construction. Overall, the controlled study shows that the evaluated private channel attacks can be monitored without training the primary monitor on attack examples and mitigated when matched counterfactual access is available.

发表机构

  • MIT Media Lab(麻省理工学院媒体实验室)
  • Westtown School(韦斯顿镇学校)
  • University of Florida(佛罗里达大学)
  • SRI International(SRI国际公司)

机构由 AI 辅助整理,请以论文原文为准。

↑