发表机构
Hanoi University of Science and Technology(河内科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出算法\Algname,将幂均值备份与多项式探索奖励结合,实现连续随机MDP中MCTS的多项式速率收敛,并实验验证其有效性。
AI 中文摘要
蒙特卡洛树搜索(MCTS)在确定性环境的在线规划中已展现出成功,但将其应用于随机马尔可夫决策过程(MDP)仍面临重大挑战,尤其是在连续状态-动作空间中。现有方法如HOOT,将MCTS与分层乐观优化(HOO)赌博机策略相结合,处理连续空间,但依赖于对数探索奖励,该奖励在非平稳随机环境中缺乏理论保证。近期进展如POLY-HOOT引入了多项式奖励项以实现确定性MDP中的收敛,但针对随机MDP的类似理论尚未发展。在本文中,我们提出了一种新颖的MCTS算法\Algname,专为连续随机MDP设计。\Algname将幂均值作为值备份算子,并结合多项式探索奖励以应对连续动作空间固有的非平稳性。我们的理论分析表明,\Algname以多项式速率$\mathcal{O}(n^{-\zeta})$收敛,其中$\zeta \in (0,1/2)$,$n$为访问的轨迹数量,从而将POLY-HOOT的非渐近收敛保证扩展到随机环境。在随机任务上的实验结果验证了我们的理论发现,展示了\Algname在连续随机领域的有效性。
英文摘要
Monte Carlo Tree Search (MCTS) has demonstrated success in online planning for deterministic environments, yet significant challenges remain in adapting it to stochastic Markov Decision Processes (MDPs), particularly in continuous state-action spaces. Existing methods, such as HOOT, which combines MCTS with the Hierarchical Optimistic Optimization (HOO) bandit strategy, address continuous spaces but rely on a logarithmic exploration bonus that lacks theoretical guarantees in non-stationary, stochastic settings. Recent advancements, such as POLY-HOOT, introduced a polynomial bonus term to achieve convergence in deterministic MDPs, though a similar theory for stochastic MDPs remains undeveloped. In this paper, we propose a novel MCTS algorithm, \Algname, designed for continuous, stochastic MDPs. \Algname integrates a power mean as a value backup operator, alongside a polynomial exploration bonus to address the non-stationarity inherent in continuous action spaces. Our theoretical analysis establishes that \Algname converges at a polynomial rate of $\mathcal{O}(n^{-ζ})$, $ζ\in (0,1/2)$, where \( n \) is the number of visited trajectories, thereby extending the non-asymptotic convergence guarantees of POLY-HOOT to stochastic environments. Experimental results on stochastic tasks validate our theoretical findings, demonstrating the effectiveness of \Algname in continuous, stochastic domains.
CommentsPublished at the International Conference on Machine Learning (ICML 2025)