arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于最优策略识别的样本高效分层强化学习

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification

Anders Jonsson, Emilie Kaufmann, Gianmarco Tedeschi, Lorenzo Steccanella

arXiv 2607.29294首次发表:更新:

AI 中文总结

该研究提出HBPI-UCRL算法,通过并行学习高低层策略,在满足特定低层动力学条件时实现多项式样本复杂度,其样本效率优于非分层算法,为分层强化学习的经验成功提供理论支撑。

AI 中文摘要

我们提出HBPI-UCRL,一种分层强化学习(HRL)的基于模型算法,可并行学习高层与低层策略。HBPI-UCRL利用高层转移对应低层多步转移的特性,对低层动力学引入两个使并行HRL可学习的充分条件。当条件满足时,我们证明HBPI-UCRL在问题参数上具有多项式样本复杂度。在稀疏奖励、目标导向场景中,HBPI-UCRL的样本复杂度上界严格低于其非分层对应算法,为HRL的经验成功提供理论依据。

英文摘要

We present HBPI-UCRL, a model-based algorithm for hierarchical reinforcement learning (HRL) that learns high-level and low-level policies in parallel. HBPI-UCRL exploits the fact that a high-level transition corresponds to a multi-step transition at the low level. We introduce two conditions on the low-level dynamics that are sufficient to make parallel HRL learnable. When these conditions hold, we prove that HBPI-UCRL has a polynomial sample complexity in the problem parameters. In the sparse-reward, goal-directed setting, our sample complexity upper bound for HBPI-UCRL is strictly lower than that of its non-hierarchical counterpart, providing theoretical justification for the empirical success of HRL.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑