发表机构
University of Oxford(牛津大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MA-JEPA提出一种基于JEPA的随机世界模型,通过预测目标表征替代观测重建,实现多智能体强化学习的集中训练与分散执行,在SMAC上表现优异。
AI 中文摘要
世界模型通过训练基于想象轨迹的策略来提高样本效率,但其效用取决于能否学习到捕获未来控制所需信息的表征。我们研究自监督联合嵌入预测(JEPA)能否为多智能体强化学习提供这种学习信号。我们提出MA-JEPA,一种随机世界模型,用目标表征的预测替代观测重建,从而支持集中训练与分散执行的多智能体基于模型的强化学习。一个分类潜在状态和一个因果Transformer通过后验与动作条件动力学预测目标进行训练,随后用于基于潜在想象的演员-评论家学习。一个仅用于训练的联合预测器以所有智能体的局部状态和动作作为条件,预测每个智能体的下一局部观测嵌入。这些预测通过与实际交互中相同的局部后验传递,并与仅用于价值学习的集中式评论家结合,执行仍保持分散。我们的实验表明,该架构在SMAC上表现强劲,在八个评估地图中的四个上,匹配或超过了所报告的最强比较器的平均胜率。
英文摘要
World models improve sample efficiency by training policies on imagined trajectories, but their usefulness depends on learning representations that capture the information needed for future control. We study whether self-supervised joint-embedding prediction (JEPA) can provide this learning signal for multi-agent reinforcement learning. We introduce MA-JEPA, a stochastic world model that replaces observation reconstruction with prediction of target representations, enabling model-based multi-agent reinforcement learning with centralized training and decentralized execution. A categorical latent state and a causal Transformer are trained with posterior and action-conditioned dynamics prediction objectives and are then used for actor-critic learning from latent imagination. A training-only joint predictor conditions on all agents' local states and actions to predict each agent's next local observation embedding. These predictions are passed through the same local posterior used during real interaction with a centralized critic that is used only for value learning, with execution remaining decentralized. Our experiments show that this architecture performs strongly on SMAC, matching or exceeding the strongest reported comparator mean win rate on four of eight evaluated maps.