发表机构
Hanoi University of Science and Technology(河内科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出基于图的随机Power-UCT算法,通过共享同深度状态提升随机MDP规划中的样本效率,并证明收敛性及在基准上的优势。
AI 中文摘要
基于树的蒙特卡洛树搜索(MCTS)在通过不同轨迹到达相同状态时会重复该状态,这在随机马尔可夫决策过程(MDP)中可能浪费模拟。我们引入了基于图的随机Power-UCT(GS-Power-UCT),它在相同规划深度到达的状态之间共享,同时为不同深度到达的状态保留独立的值。该设计适用于一般随机MDP,包括含环的问题。我们证明,对于固定规划时域,根估计以$O(n^{-1/2})$的速率收敛到有限时域值,与基于树的随机Power-UCT匹配,同时在共享状态间复用样本。我们还研究了两种全状态变体:GS-Power-UCT-F,为每个物理状态存储一个节点以增加样本共享,但可能混合不同剩余时域的值;以及GS-Power-UCT-F$^+$,使用自适应时域来控制该偏差。后者在剩余跨深度间隙消失时,收敛到根状态$s_0$的最优无限时域折扣值$V^{\star}(s_0)$。在随机规划基准上的实验表明,与基于树和基于图的基线相比,样本效率有所提高。
英文摘要
Tree-based Monte-Carlo Tree Search (MCTS) duplicates the same state when it is reached through different trajectories, which can waste simulations in stochastic MDPs. We introduce Graph-Based Stochastic-Power-UCT (GS-Power-UCT), which shares states reached at the same planning depth while keeping separate values for states reached at different depths. This design applies to general stochastic MDPs, including problems with cycles. We prove that for a fixed planning horizon, the root estimate converges to the finite-horizon value at rate $O(n^{-1/2})$, matching tree-based Stochastic-Power-UCT while reusing samples across shared states. We also study two full-state variants: GS-Power-UCT-F, which stores one node per physical state to increase sample sharing but may mix values from different remaining horizons, and GS-Power-UCT-F$^+$, which uses an adaptive horizon to control this bias. The latter converges to $V^{\star}(s_0)$, the optimal infinite-horizon discounted value at the root state $s_0$, when the remaining cross-depth gap vanishes. Experiments on stochastic planning benchmarks show improved sample efficiency over tree-based and graph-based baselines.
CommentsNo