arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.08984cs.LGcs.AIcs.GTmath.CO

稀疏奖励游戏中的AlphaZero:局限性与辅助监督

AlphaZero in Sparsely Rewarded Games: Limits and Auxiliary Supervision

Brent Kong, Tejas Ram, Tony Yue Yu

首次发表
浏览论文内容

中文总结 AI 辅助

研究在稀疏奖励游戏中普通AlphaZero的局限性,对比多帧变体及AZAL,发现普通AlphaZero虽表现出色但无法保持最优轨迹,AZAL能显著提高预言一致性,在啃咬和四子棋游戏中均有改进。

中文摘要 AI 辅助

AlphaZero已证明神经引导的蒙特卡洛树搜索可实现超人性能,但出色表现不一定意味着完美玩法。我们在两个具有对比结构的可预言评估领域进行研究:四子棋,一个具有精确博弈论值的已解决党派游戏,以及啃咬游戏,一个最优玩法由 Grundy 数结构控制的公平游戏。在统一的自我对弈+MCTS管道下,我们比较了普通AlphaZero、多帧变体(仅限于啃咬游戏)和添加预言衍生策略监督的AlphaZero辅助损失(AZAL)。我们发现普通AlphaZero在两个领域都实现了出色表现,但无法保持最优玩法所需的精确轨迹:在四子棋中,它未能保持最优玩法路线,而在啃咬游戏中,它未能始终恢复g = 0不变量。在矩形啃咬棋盘上,仅多帧输入并不能消除这一差距。然而,AZAL在多种子全游戏轨迹和采样状态评估中显著提高了预言一致性。在啃咬游戏中,AZAL在10x11上达到了完美的全游戏预言一致性,在9x10上达到了高但不完全的一致性;在四子棋中,AZAL提高了预言匹配率并延迟了第一次预言错误,但未达到完美玩法。

英文摘要

AlphaZero has demonstrated that a neural-guided Monte Carlo Tree Search can achieve superhuman performance, but strong play does not necessarily imply perfect play. We study this gap in two oracle-evaluable domains with contrasting structure: Connect Four, a solved partisan game with exact game-theoretic values, and Chomp, an impartial game whose optimal play is governed by Grundy-number structure. Under a unified self-play $+$ MCTS pipeline, we compare vanilla AlphaZero, a multi-frame variant (limited to Chomp), and an AlphaZero Auxiliary Loss (AZAL) that adds oracle-derived policy supervision. We find that vanilla AlphaZero achieves strong play across both domains but cannot preserve the exact trajectories required for optimal play: in Connect Four, it fails to maintain the optimal line of play, while in Chomp, it fails to consistently restore the $g=0$ invariant. On rectangular Chomp boards, multi-frame inputs alone do not remove this gap. Nevertheless, AZAL substantially improves oracle consistency across multi-seeded full-game traces and sampled-state evaluations. On Chomp, AZAL reaches perfect full-game oracle consistency on 10x11 and high but not complete consistency on 9x10; on Connect Four, AZAL improves oracle-match rate and delays the first oracle mistake, but does not reach perfect play.

发表机构

  • California Institute of Technology(加州理工学院)

机构由 AI 辅助整理,请以论文原文为准。

↑