arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RAVEN:面向有限字母表多智能体通信的接收者条件化动作价值编码

RAVEN: Receiver-Conditioned Action-Value Encoding for Finite-Alphabet Multi-Agent Communication

Shuwei Sun, Chenxi Wang, Jian Huang, Weiyun Ru, Hui Cao

arXiv 2609.37566首次发表:更新:

发表机构

Xi’an Jiaotong University(西安交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对有限字母表多智能体通信,提出RAVEN方法,通过接收者条件化动作价值编码训练四符号信道,在导航、捕食者-猎物及SMAC/MPE任务上以2比特消息显著提升性能。

AI 中文摘要

从一个小字母表中抽取的消息只有在保留那些会改变队友下一步决策的区分时,才能帮助队友。我们表明,按接收者情境平均的动作价值对消息进行评分,可能会恰好抹去这些区分,因此我们提出RAVEN(接收者条件化动作价值编码),它训练一个四符号、一步延迟的信道,以在每个接收者自身的上下文中保留其中心化的动作价值分布。发送者无需知道该上下文:接收者利用其私有信息解码每个符号。我们给出了该目标的两种估计器。在有教师的情况下,离线RAVEN选择能精确最小化经验条件失真的码本,并将其蒸馏到一个冻结的发送者中;我们界定了由此产生的码本选择误差和一步决策损失。在没有教师的情况下,在线RAVEN在QMIX学习器内部,将部署的符号通路与一个仅训练用的连续参考对齐,该参考共享其路由。在八个导航设置上对比五种近期通信方法,离线RAVEN在七个中取得了最高回报,而移除接收者条件化会损失其通信增益的83%。在线RAVEN将捕食者-猎物捕获成功率从53.2%提升到96.0%,相比无通信的同一QMIX骨干;在SMAC和MPE上,它在14种方法中取得了最佳平均归一化分数,包括交换千比特消息的方法。每条RAVEN消息成本为2比特,比导航上的NDQ、CACOM和ExpoComm少12至1024倍。

英文摘要

A message drawn from a small alphabet helps a teammate only if it keeps the distinctions that change that teammate's next decision. We show that scoring messages by action values averaged over the receiver's situation can erase exactly these distinctions, and we propose RAVEN (Receiver-conditioned Action-Value ENcoding), which trains a four-symbol, one-step-delayed channel to preserve each receiver's centered action-value profile within the receiver's own context. The sender never needs to know that context: the receiver decodes every symbol with its private information. We give two estimators of this target. With a teacher, offline RAVEN selects the codebook that exactly minimizes an empirical conditional distortion and distills it into a frozen sender; we bound the resulting codebook-selection error and one-step decision loss. Without a teacher, online RAVEN aligns, inside a QMIX learner, the deployed symbol pathway with a training-only continuous reference that shares its routing. Against five recent communication methods on eight navigation settings, offline RAVEN attains the highest return in seven, and removing receiver conditioning forfeits 83% of its communication gain. Online RAVEN raises predator-prey capture success from 53.2% to 96.0% over the same QMIX backbone without communication, and on SMAC and MPE it attains the best mean normalized score of 14 methods, including methods that exchange kilobit messages. Every RAVEN message costs 2 bits, 12-1,024x fewer than those of NDQ, CACOM and ExpoComm on navigation.

Comments25 pages, 20 figures, 18 tables. Code: https://github.com/sswun/RAVEN

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑