发表机构
Hanoi University of Science and Technology; Univ. Lille, Inria, CNRS, Centrale Lille, UMR 9189-CRIStAL(河内科技大学; 里尔大学、法国国家信息与自动化研究所、法国国家科学研究中心、中央理工学院里尔分校、CRIStAL实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对MCTS在模拟器与真实动态不匹配时性能下降的问题,提出鲁棒MCTS变体,通过鲁棒幂平均备份和探索奖励,实现根节点价值估计收敛速率$\mathcal{O}(n^{-1/2})$,并在显著不确定性下保持鲁棒性能。
AI 中文摘要
蒙特卡洛树搜索(MCTS)是解决复杂决策问题的强大框架,然而它常常依赖于模拟器与真实世界动态完全相同的假设。尽管这一假设有助于MCTS在国际象棋、围棋和将棋等游戏中取得成功,但在低保真模拟器中,由于建模不匹配,现实场景会引入不确定性。在这项工作中,我们提出了一种新的鲁棒MCTS变体,以缓解动态模型的不确定性。我们的算法处理转移动态和奖励分布的不确定性,以弥合基于模拟的规划与现实部署之间的差距。我们整合了鲁棒幂平均备份算子以及精心设计的探索奖励,以确保搜索树中每个节点的有限样本收敛性。我们证明了我们的算法在根节点的价值估计上实现了$\mathcal{O}(n^{-1/2})$的收敛速率,与标准MCTS相当。最后,我们提供了实证证据,表明即使在底层奖励分布和转移动态存在显著不确定性的情况下,我们的方法在规划问题中也能实现鲁棒性能。
英文摘要
Monte Carlo Tree Search (MCTS) is a powerful framework for solving complex decision-making problems, yet it often relies on the assumption that the simulator and the real-world dynamics are identical. Although this assumption helps achieve the success of MCTS in games like Chess, Go, and Shogi, the real-world scenarios incur ambiguity due to their modeling mismatches in low-fidelity simulators. In this work, we present a new robust variant of MCTS that mitigates dynamical model ambiguities. Our algorithm addresses transition dynamics and reward distribution ambiguities to bridge the gap between simulation-based planning and real-world deployment. We incorporate a robust power mean backup operator and carefully designed exploration bonuses to ensure finite-sample convergence at every node in the search tree. We show that our algorithm achieves a convergence rate of $\mathcal{O}(n^{-1/2})$ for the value estimation at the root node, comparable to that of standard MCTS. Finally, we provide empirical evidence that our method achieves robust performance in planning problems even under significant ambiguity in the underlying reward distribution and transition dynamics.