PureTD:无评估时间搜索的西洋双陆金钱游戏强化学习
PureTD: Reinforcement Learning for Backgammon Money Games with No Evaluation-time Search
AI总结:
该研究提出PureTD模型,通过纯自玩强化学习在无评估时间搜索下训练西洋双陆金钱游戏,其无搜索模型比两款开源引擎更强且评估速度更快。
AI中文摘要:
我们在无评估时间搜索的场景下重新研究Tesauro的TD-Gammon西洋双陆金钱游戏,通过自玩强化学习(RL)从头学习棋子走法和加倍骰子使用(cube action),仅保留极少手工编码逻辑且无专家特征。在该场景下,我们证明纯自玩RL足以训练出达到接近当前顶尖水平的模型。具体而言,对于含加倍骰子的金钱游戏,我们的无搜索模型评估速度更快,且比运行1步(1-ply)前瞻搜索的开源引擎GNU Backgammon和Open Sage强得多。
英文摘要:
We revisit Tesauro's TD-Gammon for backgammon money games in the setting of no evaluation-time search. Both checker play and cube action (use of the doubling cube) are learned from scratch via self-play reinforcement learning (RL), with minimal hand-coded logic and no expert features. In this setting, we demonstrate that pure self-play RL suffices to train models that reach near-state-of-the-art playing strength. Specifically, for cubeful money games, our search-free model evaluates faster and is substantially stronger than the open-source engines GNU Backgammon and Open Sage running a one-move (1-ply) look-ahead search.