arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

NashDreamer:用于零和非完全信息博弈的基于模型的强化学习

NashDreamer: Model-Based Reinforcement Learning for Zero-Sum Imperfect-Information Games

Tomáš Holeček, Viliam Lisý

arXiv 2609.01549首次发表:更新:

发表机构

Artificial Intelligence Center; Faculty of Electrical Engineering, Czech Technical University in Prague(人工智能中心; 布拉格捷克技术大学电气工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对基于模型的强化学习在零和非完全信息博弈中应用不足的问题,提出NashDreamer框架,引入MARSSM模型,经四基准游戏验证可提升样本效率,并分析了Dreamer系列算法的后验崩溃问题。

AI 中文摘要

基于模型的强化学习(MBRL)在单智能体领域已取得显著成果,但其向竞争性非完全信息博弈(IIGs)的扩展仍未得到充分探索。在多智能体场景中,对手引发的非平稳性使学习过程复杂化,而分散式模型学习面临严重的可识别性障碍,我们认为这使得集中式模型学习成为数学上的必然要求。基于该分析,我们提出NashDreamer,一种用于两人零和IIGs的原则性MBRL框架。它引入了集中式多智能体循环状态空间模型(MARSSM),该模型将环境动态与玩家策略对其各自观测的影响解耦。NashDreamer设计用于使用任意策略梯度算法,并在理想模型下继承它们向纳什均衡收敛的保证。对四个基准游戏的实证评估表明,NashDreamer在训练早期比无模型基线显著提高了样本效率。最后,我们从理论上分析了该架构的优化景观,确定了Dreamer系列算法在随机环境中对后验崩溃的脆弱性,我们将此留作一个开放挑战。

英文摘要

Model-based reinforcement learning (MBRL) has achieved remarkable results in single-agent domains, yet its extension to competitive imperfect information games (IIGs) remains underexplored. In multi-agent settings, opponent-induced non-stationarity complicates the learning process, and decentralized model learning faces severe identifiability barriers, which we argue make centralized model learning a mathematical necessity. Building on this analysis, we propose NashDreamer, a principled MBRL framework for two-player zero-sum IIGs. It introduces a centralized Multi-Agent Recurrent State-Space Model (MARSSM) that decouples environment dynamics from the effect of players' strategies on their individual observations. NashDreamer is designed to use arbitrary policy gradient algorithms and inherits their convergence guarantees towards Nash equilibria under an idealized model. Empirical evaluations across four benchmark games demonstrate that NashDreamer substantially improves sample efficiency over model-free baselines early in the training. Finally, we theoretically analyze the architecture's optimization landscape, identifying the vulnerability of the Dreamer family of algorithms to posterior collapse in stochastic environments. We leave it as an open challenge.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑