发表机构
Czech Technical University in Prague; Carnegie Mellon University; Artificial Intelligence Center; Strategy Robot, Inc.; Strategic Machine, Inc.; Optimized Markets, Inc.(布拉格捷克理工大学; 卡内基梅隆大学; 人工智能中心; 策略机器人公司; 战略机器公司; 优化市场公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出一种结合磁镜下降与高斯混合重参数化的策略梯度算法,用于连续动作博弈,通过自我对弈训练,在多个博弈中优于现有方法并减少样本需求。
AI 中文摘要
大多数超人级游戏算法的成功都集中在离散动作的游戏中,然而在拍卖、机器人、体育或交易等领域,动作几乎是连续的。现有技术要么依赖专家设计的离散化,要么样本效率低下。我们提出了一种可扩展的策略梯度算法,用于具有连续或混合离散与连续动作的大型序贯博弈。该算法将磁镜下降与高斯混合重参数化相结合,并通过自我对弈进行训练。我们证明,在梯度下降失效的博弈中,该算法能够逼近均衡。在序贯博弈中,它优于神经虚拟自我对弈,并以少3.5至5.5倍的样本达到或超越策略空间响应预言机的最终策略。在单挑无限注德州扑克中,其表现与Slumbot相当。
英文摘要
Most successes of superhuman game-playing algorithms are in games with discrete actions, yet in auctions, robotics, sports, or trading, actions are nearly continuous. Prior techniques either rely on expert-designed discretizations or are sample inefficient. We present a scalable policy-gradient algorithm for large sequential games with continuous or mixed discrete and continuous actions. It combines magnetic mirror descent with a mixture of Gaussians reparametrization, trained via self-play. We show that it approximates equilibrium in games where gradient descent fails. In sequential games, it outperforms neural fictitious self-play and matches or outperforms the final strategies of policy space response oracles with 3.5--5.5$\times$ fewer samples. In heads-up no-limit Texas hold'em, it performs on par with Slumbot.