涌现,而非带宽:物理耦合与学习型多智能体通信的极限
Emergence, Not Bandwidth: Physical Coupling and the Limits of Learned Multi-Agent Communication
浏览论文内容
中文总结 AI 辅助
本研究证明在速率受限多智能体系统中,通信价值由物理耦合决定而非带宽,并揭示强化学习在部分耦合下无法发现最优协议,导致性能显著低于工程化方案。
中文摘要 AI 辅助
速率受限的多智能体团队提出了三个问题,而涌现通信文献仅凭经验回答了这些问题:最优消息应编码什么、压缩在时间跨度上的代价是什么,以及学习到的协议何时足够独特以便队友解读。我们针对速率受限的Dec-POMDP回答了这些问题,然后衡量强化学习与最优解之间的差距。我们的定理确定了与任何学习者无关的可实现性,因此在相同比特预算下,工程化发送器与学习型发送器之间的差距是优化问题,而非信息论问题。我们在三个MuJoCo竞技场上进行了实例化,涵盖零耦合、部分耦合和刚性物理耦合,每个决策恰好分配2比特,并通过在一个竞技场内关闭物理侧信道来创建判别性条件,保持物体、任务和奖励不变。通信价值由耦合决定:在通过共享物体的刚性耦合下,任何信道都不优于沉默(+0.001 +/- 0.001,p = 0.982,n = 25),因为本体感觉已携带该信息;无耦合时,所有条件均能解决任务;部分耦合下,工程化2比特发送器达到四分位均值1.000,而学习型发送器仅达0.482,与沉默无法区分(p = 0.400,n = 25)。在共享字母表下,带宽无法解释这一差距。从工程化接收器热启动定位了失败:同一信道热启动达到0.857,而冷启动为0.562(p < 0.001),因此既非表征问题也非维护问题;强化学习未能发现该协议。交叉对局显示学习到的协议个体有意义但相互不可理解:自对局0.980在不同种子间降至0.144,而我们最佳构建的对齐仍留下至少77%的差距。所有主要结果均使用每个竞技场25个种子和七个已发布基线在匹配速率下获得。
英文摘要
Rate-limited multi-agent teams raise three questions the emergent-communication literature has answered only empirically: what an optimal message should encode, what compression costs over a horizon, and when a learned protocol is unique enough for a teammate to read. We answer them for rate-limited Dec-POMDPs, then measure how far reinforcement learning falls short of the optimum. Our theorems fix what is achievable independently of any learner, so a gap between an engineered and a learned sender at the same bit budget is an optimization fact, not an information-theoretic one. We instantiate this on three MuJoCo arenas spanning zero, partial and rigid physical coupling, charging every condition exactly 2 bits per decision, and create the discriminating regime by closing a physical side channel within one arena, holding bodies, task and reward fixed. Communication value is governed by coupling: under rigid coupling through a shared object, no channel beats silence (+0.001 +/- 0.001, p = 0.982, n = 25), since proprioception already carries that information; without coupling, every condition solves the task; under partial coupling, the engineered 2-bit sender reaches an interquartile mean of 1.000 but the learned one reaches 0.482, indistinguishable from silence (p = 0.400, n = 25). With a shared alphabet, bandwidth cannot explain the gap. Warm-starting from an engineered receiver localizes the failure: the same channel reaches 0.857 versus 0.562 cold-started (p < 0.001), so it is neither representational nor one of maintenance; reinforcement learning fails to discover the protocol. Cross-play shows learned protocols are individually meaningful but mutually unintelligible: self-play 0.980 collapses to 0.144 across seeds, and our best constructed alignment leaves at least 77% of that gap. All headline results use 25 seeds per arena and seven published baselines at matched rate.
发表机构
- Purdue University(普渡大学)
机构由 AI 辅助整理,请以论文原文为准。