arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

正则化策略梯度与学习的高斯混合用于连续动作博弈

Regularized policy gradient with learned mixtures of Gaussians for games with continuous actions

Ondřej Kubíček, Viliam Lisý, Tuomas Sandholm

arXiv 2609.36787首次发表:更新:

发表机构

Czech Technical University in Prague; Carnegie Mellon University; Artificial Intelligence Center; Strategy Robot, Inc.; Strategic Machine, Inc.; Optimized Markets, Inc.(布拉格捷克理工大学; 卡内基梅隆大学; 人工智能中心; 策略机器人公司; 战略机器公司; 优化市场公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出一种结合磁镜下降与高斯混合重参数化的策略梯度算法,用于连续动作博弈,通过自我对弈训练,在多个博弈中优于现有方法并减少样本需求。

AI 中文摘要

大多数超人级游戏算法的成功都集中在离散动作的游戏中,然而在拍卖、机器人、体育或交易等领域,动作几乎是连续的。现有技术要么依赖专家设计的离散化,要么样本效率低下。我们提出了一种可扩展的策略梯度算法,用于具有连续或混合离散与连续动作的大型序贯博弈。该算法将磁镜下降与高斯混合重参数化相结合,并通过自我对弈进行训练。我们证明,在梯度下降失效的博弈中,该算法能够逼近均衡。在序贯博弈中,它优于神经虚拟自我对弈,并以少3.5至5.5倍的样本达到或超越策略空间响应预言机的最终策略。在单挑无限注德州扑克中,其表现与Slumbot相当。

英文摘要

Most successes of superhuman game-playing algorithms are in games with discrete actions, yet in auctions, robotics, sports, or trading, actions are nearly continuous. Prior techniques either rely on expert-designed discretizations or are sample inefficient. We present a scalable policy-gradient algorithm for large sequential games with continuous or mixed discrete and continuous actions. It combines magnetic mirror descent with a mixture of Gaussians reparametrization, trained via self-play. We show that it approximates equilibrium in games where gradient descent fails. In sequential games, it outperforms neural fictitious self-play and matches or outperforms the final strategies of policy space response oracles with 3.5--5.5$\times$ fewer samples. In heads-up no-limit Texas hold'em, it performs on par with Slumbot.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑