arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23269cs.ROcs.LGcs.MA

潜在心灵感应:基于自监督感知潜在向量的多机器人通信

Latent Telepathy: Multi-Robot Communication with Self-Supervised Perceptual Latents

Howard Wang, Han Zheng, Cathy Wu

首次发表
浏览论文内容

中文总结 AI 辅助

提出潜在心灵感应方法,通过广播自监督编码器生成的感知潜在向量实现多机器人通信,在带宽受限下以99.7%成功率避开遮挡危险,优于位置、轨迹和原始图像消息。

中文摘要 AI 辅助

在部分可观测条件下的分散式多机器人团队中,决定机器人下一步行动的事实往往仅对队友可见。现有的分散式方法通信运动学信息(如位置或规划轨迹),无法传达队友的感知内容。多智能体强化学习(MARL)中的学习通信可以携带感知内容,但产生的消息与任务耦合且不透明。我们提出潜在心灵感应(Latent Telepathy)。每个机器人广播其已为自身使用而计算的感知潜在向量,该向量是使用自监督联合嵌入预测目标训练的编码器的输出,该编码器被冻结并在团队中共享。队友仅从任务奖励中学习基于该向量行动。由于编码器已用于感知,消息不增加额外计算成本,仅需一个紧凑的带宽向量。由于编码器在任何策略训练前被冻结,消息对每个机器人含义相同,接收机器人从未被告知其含义。我们通过内容受控协议评估潜在心灵感应,其中带宽、延迟、拓扑和接收器固定,仅消息内容变化。广播潜在向量使导航器在99.7%的回合中避开遮挡危险,与无噪声手工设计消息相当。位置和轨迹消息仍处于随机水平,而原始相机图像(带宽宽186倍)比压缩潜在向量可靠性更低。该结果从离散网格世界到连续速度控制下的渲染像素均成立,编码器在102次实时决策中的102次从物理机器人相机解码出危险。我们还确定了将MARL通信结果移植到连续控制的一个要求,即消息所告知的决策必须可通过探索到达,并展示了如何恢复该要求。

英文摘要

In a decentralized multi-robot team under partial observability, the fact that decides a robot's next action is often visible only to a teammate. Existing decentralized methods communicate kinematic information, such as position or planned trajectory, which cannot convey what the teammate perceives. Learned communication in multi-agent reinforcement learning (MARL) can carry perceptual content, but the resulting messages are task-coupled and opaque. We propose Latent Telepathy. Each robot broadcasts the perceptual latent vector it already computes for its own use, the output of an encoder trained with a self-supervised joint-embedding predictive objective, frozen, and shared across the team. A teammate learns to act on it from task reward alone. Because the encoder already runs for perception, the message costs no additional computation and a single compact vector of bandwidth. Because the encoder is frozen before any policy is trained, the message means the same thing to every robot, and the receiving robot is never told what it means. We evaluate Latent Telepathy with a content-controlled protocol in which bandwidth, latency, topology and receiver are held fixed and only the message content varies. Broadcasting the latent lets a navigator avoid an occluded hazard in 99.7% of episodes, matching a noiseless hand-designed message. Position and trajectory messages remain at chance, and the raw camera image, 186 times wider, is less reliable than the compressed latent. The result holds from a discrete gridworld to rendered pixels under continuous velocity control, and the encoder decodes the hazard from a physical robot's camera in 102 of 102 live decisions. We also identify a requirement for porting MARL communication results to continuous control, that the decision a message informs must remain reachable by exploration, and show how to restore it.

发表机构

  • Columbia University(哥伦比亚大学)
  • Massachusetts Institute of Technology(麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑