面向持续分层强化学习与规划的自主选项发明
Autonomous Option Invention for Continual Hierarchical Reinforcement Learning and Planning
- Arizona State University(亚利桑那州立大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出在持续强化学习中自主发明并利用具有符号表示的高级选项的方法,通过结合状态抽象与前瞻规划,实现了选项的可组合、可重用与互独立,显著提升了跨任务的样本效率与知识迁移能力。
AI中文摘要:
抽象是扩大强化学习(RL)规模的关键。然而,自主学习抽象的状态和动作表示以实现迁移和泛化,仍是一个极具挑战性的开放问题。本文提出了一种在持续RL环境中发明、表示和利用选项(代表时间上扩展的行为)的新方法。我们的方法处理具有长视野、稀疏奖励以及未知转移和奖励函数的随机问题流。该方法持续学习并维护可解释的状态抽象,并利用它来发明具有抽象符号表示的高级选项。这些选项满足三个关键要求:(1)可组合性,通过前瞻规划有效解决任务;(2)可重用性,跨问题实例重用以最小化重新学习的需求;(3)互独立性,减少选项间的干扰。我们的主要贡献是持续学习具有符号表示的可迁移、可泛化选项的方法,以及将搜索技术与RL集成以在这些已学选项上进行高效规划以解决新问题的方法。实验结果表明,该方法能有效地跨问题实例学习和迁移抽象知识,与最先进的方法相比实现了更优的样本效率。
英文摘要:
Abstraction is key to scaling up reinforcement learning (RL). However, autonomously learning abstract state and action representations to enable transfer and generalization remains a challenging open problem. This paper presents a novel approach for inventing, representing, and utilizing options, which represent temporally extended behaviors, in continual RL settings. Our approach addresses streams of stochastic problems characterized by long horizons, sparse rewards, and unknown transition and reward functions. Our approach continually learns and maintains an interpretable state abstraction, and uses it to invent high-level options with abstract symbolic representations. These options meet three key desiderata: (1) composability for solving tasks effectively with lookahead planning, (2) reusability across problem instances for minimizing the need for relearning, and (3) mutual independence for reducing interference among options. Our main contributions are approaches for continually learning transferable, generalizable options with symbolic representations, and for integrating search techniques with RL to efficiently plan over these learned options to solve new problems. Empirical results demonstrate that the resulting approach effectively learns and transfers abstract knowledge across problem instances, achieving superior sample efficiency compared to state-of-the-art methods.