BluffJAX:JAX中的对抗性不完美信息博弈
BluffJAX: Adversarial Imperfect Information Games in JAX
浏览论文内容
中文总结 AI 辅助
BluffJAX是一个基于JAX的开源套件,提供多种对抗性不完美信息博弈(如德州扑克、Kuhn扑克等)的高吞吐量GPU并行实现,并给出基准测试结果,以促进博弈论强化学习研究。
中文摘要 AI 辅助
我们介绍了BluffJAX:一个基于JAX的开源对抗性不完美信息博弈套件。我们提供了专为高模拟吞吐量和GPU加速器并行化设计的博弈的规范实现。我们的套件包含诸如德州扑克和Kuhn扑克等研究充分的基准,以及此前未在强化学习研究中研究过的博弈,如Bluff、Stud扑克和Kemps。我们希望实现多种博弈机制和难度,将为强化学习的博弈论方法引入新的挑战并促进新的研究方向。我们在单GPU和多GPU设置下对我们的环境的吞吐量性能和内存使用进行了基准测试,展示了每秒高达数亿样本的扩展能力,并激励使用BluffJAX而非相关的基于GPU和CPU的库。我们在JAX中对强化学习、树搜索和博弈求解算法进行基准测试,以便为用户提供基线结果并促进未来的比较。
英文摘要
We introduce BluffJAX: an open-source suite of adversarial imperfect information games in JAX. We provide canonical implementations of games designed for high simulation throughputs and parallelization on GPU accelerators. Our suite consists of well-studied benchmarks such as Texas Hold'Em Poker and Kuhn Poker, as well as games that have not been previously studied in reinforcement learning research, such as Bluff, Stud Poker, and Kemps. We hope that implementing a variety of game mechanics and difficulties will introduce new challenges and foster novel research directions in game-theoretic methods for RL. We benchmark the throughput performance and memory usage of our environments in single and multi-GPU settings, demonstrating scaling of up to hundreds of millions of samples per second, and motivating the usage of BluffJAX over related GPU and CPU-based libraries. We benchmark reinforcement learning, tree search, and game-solving algorithms in JAX in order to provide users with baseline results and facilitate future comparisons.
发表机构
- TU Darmstadt(达姆施塔特工业大学)
- Hessian Center for Artificial Intelligence (Hessian.ai)(黑森人工智能中心)
- German Research Center for AI (DFKI)(德国人工智能研究中心)
- University of Würzburg(维尔茨堡大学)
机构由 AI 辅助整理,请以论文原文为准。