arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

蒙特卡洛树搜索只是全访问蒙特卡洛控制吗?

Is Monte Carlo Tree Search Just Every-Visit Monte Carlo Control?

Xianyi Wu

arXiv 2608.27985首次发表:更新:

发表机构

ECNU(华东师范大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文论证蒙特卡洛树搜索(MCTS)与全访问蒙特卡洛控制本质等价,仅表述术语不同,MCTS可简化为轨迹采样与全访问蒙特卡洛更新两个操作,旨在明确二者的等价关系。

AI 中文摘要

蒙特卡洛树搜索(MCTS)和全访问蒙特卡洛(MC)控制通常被视为不同的方法:MCTS以搜索语言(选择、扩展、模拟、回溯)描述,MC控制以强化学习语言(轨迹采样、回报估计、动作值更新、策略改进)描述。本文指出,在轨迹生成和动作值更新层面,二者的差异主要是术语层面的:树策略和展开策略可视为单一演化策略中已学习和未学习的部分;扩展对应首次访问与初始化;回溯是常规的全访问蒙特卡洛更新。按此解读,MCTS的四个阶段可简化为两个基本操作:当前策略下的轨迹采样与全访问蒙特卡洛更新。从这个意义上讲,MCTS只是用搜索的语言和数据结构表述的全访问蒙特卡洛控制。本文的目的是说明这一等价关系,使其更易被识别。

英文摘要

Monte Carlo Tree Search (MCTS) and every-visit Monte Carlo (MC) control are usually presented as different methods. MCTS is described in the language of search (selection, expansion, simulation, and backup), whereas MC control is described in the language of reinforcement learning (trajectory sampling, return estimation, action-value updating, and policy improvement). This note argues that, at the level of trajectory generation and action-value updating, the distinction is largely terminological. The tree policy and rollout policy can be viewed as the learned and not-yet-learned parts of a single evolving policy; expansion corresponds to first visit and initialization; and backup is the ordinary every-visit Monte Carlo update. Under this interpretation, the four stages of MCTS reduce to two basic operations: trajectory sampling under the current policy and every-visit Monte Carlo updating. In this sense, MCTS is simply every-visit Monte Carlo control expressed in the language and data structure of search. The purpose of this note is expository: to make this equivalence explicit and easier to recognize.

CommentsComments and discussions are welcome

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑