arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.19809cs.MAcs.LG

Dreamer-CPC:用于去中心化多智能体强化学习的基于世界模型的消息学习

Dreamer-CPC: Message Learning with World Models for Decentralized Multi-agent Reinforcement Learning

Taisuke Takayama, Naoto Yoshida, Tadahiro Taniguchi

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对多智能体强化学习中通信问题,提出Dreamer-CPC方法,将基于CPC的消息学习集成到DreamerV3世界模型,各智能体据此独立维护模型与模块并交换消息,实验表明该方法在多种环境下优于现有方法,能有效支持去中心化决策。

中文摘要 AI 辅助

在多智能体强化学习(MARL)中,智能体间通信对于在部分可观测性下提高性能很有效。基于表征学习的方法能让去中心化智能体基于自身观测学习消息,但仅依赖当前观测,无法传递随时间积累的信息。我们提出Dreamer-CPC,一种基于模型的去中心化MARL方法,将基于集体预测编码(CPC)的消息学习集成到DreamerV3的世界模型中。每个智能体独立维护一个世界模型和一个消息模块,并从反映过去观测和行动历史的世界模型的潜在状态中推断和交换消息。我们在Observer(一个非合作信息共享任务)和CatchApple(一个新引入的任务,其中与任务相关的观测暂时缺失)这两个环境中评估了Dreamer-CPC。在这两个环境中,Dreamer-CPC均优于IPPO-CPC(一种现有的基于CPC的方法,从当前观测生成消息)以及无通信基线。特别是在CatchApple中,Dreamer-CPC实现的回合回报是IPPO-CPC的4到5倍,展示了在其他方法因观测缺失而失败时的有效协调。这些结果表明,当仅当前观测不足时,基于世界模型潜在动态的通信可支持去中心化决策。

英文摘要

In multi-agent reinforcement learning (MARL), inter-agent communication is effective for improving performance under partial observability. Representation learning-based approaches enable decentralized agents to learn messages grounded in their own observations, but they rely only on current observations and cannot convey information accumulated over time. We propose Dreamer-CPC, a decentralized model-based MARL method that integrates message learning based on Collective Predictive Coding (CPC) into the world model of DreamerV3. Each agent independently maintains a world model and a message module, and infers and exchanges messages from the latent states of the world model that reflect the history of past observations and actions. We evaluated Dreamer-CPC in two environments: Observer, a non-cooperative information-sharing task, and CatchApple, a newly introduced task in which task-relevant observations are temporarily missing. In both environments, Dreamer-CPC outperformed IPPO-CPC, an existing CPC-based method that generates messages from current observations, as well as no-communication baselines. In particular, in CatchApple, Dreamer-CPC achieved 4 to 5 times the episode return of IPPO-CPC, demonstrating effective coordination where other methods fail due to missing observations. These results suggest that communication grounded in the latent dynamics of world models can support decentralized decision-making when current observations alone are insufficient.

发表机构

  • Graduate School of Informatics, Kyoto University, Kyoto, Japan(京都大学信息学研究生院)
  • Research Organization of Science and Technology, Ritsumeikan University, Shiga, Japan(立命馆大学科学技术研究机构)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑