AI 中文总结
研究Baghchal非对称棋盘游戏,用深度Q网络、REINFORCE、近端策略优化和MuZero四种方法训练并评估,比较各算法在胜率等方面表现,发现MuZero性能最佳,PPO实用且计算成本低,揭示不同策略在游戏中的特点。
AI 中文摘要
Baghchal是一款起源于尼泊尔的两人非对称棋盘游戏,四只老虎捕捉二十只山羊。其结构复杂、具策略性、信息完备且有文化意义,但在深度强化学习文献中未被充分研究。本文系统探索了深度Q网络(DQN)、REINFORCE、近端策略优化(PPO)和MuZero这四种深度强化学习解决方案,在Baghchal非对称游戏的一侧训练,另一侧评估。基于胜率、平局率、平均捕获数、训练收敛和计算成本对算法进行评级。实验发现MuZero在两项任务中性能最佳,PPO是最实用的算法,计算成本显著低于MuZero。紧急战略行为分析表明,基于模型的策略在长期规划中最优,而像DQN这样基于价值的策略因奖励信号更强而更偏向老虎角色。
英文摘要
Baghchal is a two-player asymmetric board game with Nepali origins where four tigers are to capture goats and twenty goats desire to keep tigers in immobility. Although Baghchal has a complex structure which is strategic, has perfect information structure, and has cultural meaning, it has not been adequately covered in deep reinforcement learning (RL) literature. This paper gives a systematic exploration of four deep RL solutions Deep Q-Network (DQN), REINFORCE, Proximal Policy Optimization (PPO) and MuZero that are trained on one side of the asymmetric gameplay of Baghchal and then evaluated on the other side. The algorithms are rated based on win rate, draw rate, average captures, training convergence and computational cost. It is experimentally found that MuZero generates the best performance in both tasks, achieving 86 percent win over these Tiger and 62 percent win over these Goat and the ability to do so is due to the model-based planning machine through the Monte Carlo Tree Search. PPO is the most realistic algorithm and is provided to be competitive over both asymmetric tasks with significantly reduced computational costs compared to MuZero. Emergent strategic behavior analysis shows that model-based strategies are optimal over long-horizon planning, whereas value-based counterparts like DQN are more biased up towards the Tiger role owing to the more substantial reward signal.
Comments6 pages, 6 figures