发表机构
Sogang University(西江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对模型未知的有限时域MDP,将自适应多阶段采样转化为无采样算法AMR,在强化学习环境中实现渐近最优的值估计。
AI 中文摘要
本文提出了一种学习方法,用于在给定有限马尔可夫决策过程(MDP)的底层模型对决策者未知的情况下,求解有限时域马尔可夫决策过程。我们将自适应多阶段采样(AMS)算法转换为一种无采样算法,称为“自适应多阶段滚动(AMR)”,用于在仅知道状态集和动作集时估计初始状态的最优值。AMR 模拟了 AMS 中的反向归纳,但处于强化学习(RL)环境中。在每次迭代中,AMR 生成一个非平稳策略用于探索,并滚动执行该策略以获得单一的经验轨迹,然后以非递归方式反向追踪该轨迹,仅在访问过的状态和采取过的动作处进行相关更新。我们证明 AMR 是渐近最优的,即期望绝对误差序列趋近于零,其收敛速度取决于从初始状态出发在每个阶段对每个可达状态的访问次数,从而将 AMS 的结果本质性地转化到 RL 环境中。
英文摘要
This work provides a learning approach to solving finite-horizon Markov decision processes (MDPs) when the underlying model of a given finite MDP is unknown to the decision maker. We transform the adaptive multistage sampling (AMS) algorithm into a sampling-free algorithm, called "adaptive multistage rollout (AMR)," for estimating the optimal value at an initial state when only the state set and the action set are known. AMR emulates the backward induction as in AMS but in a reinforcement learning (RL) setting. At each iteration, AMR generates a non-stationary policy to be used for exploration and rolls out the policy in order to obtain a single trajectory of experiences and traces it backwards in a non-recursive way while doing relevant updates only at visited states and for actions taken at the visited states. We show that AMR is asymptotically optimal such that the sequence of the expected absolute errors approaches zero and its convergence rate depends on the number of visits to each reachable state at each stage from the initial state, essentially transforming the result of AMS into the RL setting.