arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.14555eess.SYcs.SY

有限时域马尔可夫决策过程树搜索中Q函数估计器的一致方差估计

Consistent Variance Estimation for Q-Function Estimators in Finite-Horizon MDP Tree Search

Zhenyu Yue, Jie Xu, Chun-Hung Chen, Hadi El-Amine, Michael C. Fu

首次发表
浏览论文内容

中文总结 AI 辅助

研究有限时域MDP树搜索中Q函数估计器方差,指出基于独立同分布路径假设的估计器有偏差,提出一致的递归方差估计器并给出等效实现,集成到MCTS采样过程中,在数值示例中提升了MCTS算法性能。

中文摘要 AI 辅助

我们研究了有限时域、有限状态马尔可夫决策过程(MDP)树搜索中Q函数估计器的方差。方差分解为即时奖励、概率状态转移和未来状态值函数估计的不确定性三个部分。基于独立同分布路径假设的样本方差估计器有偏差且低估真实方差。我们提出了一个一致的递归方差估计器,并给出仅使用可迭代更新的节点局部统计量的等效实现。该估计器被集成到有限时域MDP的两种蒙特卡洛树搜索(MCTS)采样过程中,在库存控制和肾脏配对捐赠匹配的数值示例中,新估计器相对于使用基于独立同分布样本方差估计器的基线提高了MCTS算法的性能。

英文摘要

We study the variance of Q-function estimators in finite-horizon, finite-state Markov decision process (MDP) tree search. We show that the variance decomposes into three components attributed to the immediate reward collected, probabilistic state transitions, and uncertainty in future state value function estimates. Using this decomposition, we show that the sample variance estimator based on the assumption of i.i.d. paths is biased, underestimating the true variance, and the bias does not vanish in the limit. We then propose a recursive variance estimator that is consistent. To enable efficient storage and computation, we derive an equivalent implementation of the recursive estimator using only node-local statistics that can be iteratively updated. This consistent variance estimator is integrated into two Monte Carlo Tree Search (MCTS) sampling procedures for finite-horizon MDPs. In numerical examples from inventory control and kidney paired donation matching, the new estimator improves the performance of the MCTS algorithm relative to a baseline that uses the i.i.d.-based sample variance estimator.

↑