arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15146cs.LG

PureTD:无评估时间搜索的西洋双陆金钱游戏强化学习

PureTD: Reinforcement Learning for Backgammon Money Games with No Evaluation-time Search

Alexander L. Strehl

AI总结:

该研究提出PureTD模型,通过纯自玩强化学习在无评估时间搜索下训练西洋双陆金钱游戏,其无搜索模型比两款开源引擎更强且评估速度更快。

AI中文摘要:

我们在无评估时间搜索的场景下重新研究Tesauro的TD-Gammon西洋双陆金钱游戏,通过自玩强化学习(RL)从头学习棋子走法和加倍骰子使用(cube action),仅保留极少手工编码逻辑且无专家特征。在该场景下,我们证明纯自玩RL足以训练出达到接近当前顶尖水平的模型。具体而言,对于含加倍骰子的金钱游戏,我们的无搜索模型评估速度更快,且比运行1步(1-ply)前瞻搜索的开源引擎GNU Backgammon和Open Sage强得多。

英文摘要:

We revisit Tesauro's TD-Gammon for backgammon money games in the setting of no evaluation-time search. Both checker play and cube action (use of the doubling cube) are learned from scratch via self-play reinforcement learning (RL), with minimal hand-coded logic and no expert features. In this setting, we demonstrate that pure self-play RL suffices to train models that reach near-state-of-the-art playing strength. Specifically, for cubeful money games, our search-free model evaluates faster and is substantially stronger than the open-source engines GNU Backgammon and Open Sage running a one-move (1-ply) look-ahead search.

↑