arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多链马尔可夫决策过程的平均奖励强化学习:一种层次分解方法

Average-Reward Reinforcement Learning for Multichain MDPs: A Hierarchical Decomposition Approach

Huizhen Yu, Isaiah Heidt

arXiv 2610.10326首次发表:更新:

发表机构

University of Alberta(阿尔伯塔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多链MDP的平均奖励强化学习,提出基于Bather分解的层次化异步值迭代算法,收敛到最优增益并生成增益最优策略,进一步扩展两种算法提升瞬态性能,为一般多链MDP提供首批无模型平均奖励RL方法。

AI 中文摘要

我们研究平均奖励多链马尔可夫决策过程(MDPs)中的最优策略学习问题,其中最优增益可能依赖于初始状态,且不同策略下的递归结构各异,这给强化学习(RL)方法带来了挑战。我们提出了一种基于异步值迭代的RL算法,该算法除了MDP的转移图外不需要任何模型知识,并利用Bather分解将状态空间层次化地划分为通信子系统和瞬态状态。这种分解将全局决策问题重构为结构化的子问题,我们的算法充分利用了这一点。我们证明该算法收敛到最优增益,并在有限时间后产生增益最优策略。在此基础算法之上,我们进一步开发了两种算法:一种近似求解多链平均最优性方程以获得接近增益最优的策略,另一种通过近似最优偏差函数并使用基础算法求解一个诱导的平均奖励多链MDP,以实现接近偏差最优性。我们为所有三种算法提供了几乎必然收敛的保证,并实证比较了它们的权衡,表明后两种算法在瞬态性能上也持续优于基础算法。据我们所知,这是首批无需模型知识、不通过折扣问题归约而适用于一般多链MDP的平均奖励RL算法。

英文摘要

We study learning optimal policies in average-reward multichain Markov decision processes (MDPs), where the optimal gain may depend on the initial state and recurrence structures vary across policies, creating challenges for reinforcement learning (RL) methods. We propose an asynchronous value-iteration-based RL algorithm that requires no model knowledge beyond the MDP's transition graph and leverages Bather's decomposition to hierarchically partition the state space into communicating subsystems and transient states. This decomposition induces a recasting of the global decision problem into structured subproblems, which our algorithm exploits. We show that the algorithm converges to the optimal gain and produces gain-optimal policies after finite time. Building on this base algorithm, we develop two further algorithms: one approximately solves the multichain average optimality equations to obtain near gain-optimal policies, and another targets near bias-optimality by approximating the optimal bias function and solving an induced average-reward multichain MDP using the base algorithm. We provide almost-sure convergence guarantees for all three algorithms and empirically compare their tradeoffs, showing that the latter two also consistently improve transient performance relative to the base algorithm. To our knowledge, these are the first essentially model-free average-reward RL algorithms for general multichain MDPs without reductions to discounted problems.

Comments60 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑