arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2112.09836cs.AIcs.LG

人工智能的创造力:促进深度强化学习的分层规划模型学习

Creativity of AI: Hierarchical Planning Model Learning for Facilitating Deep Reinforcement Learning

  • Sun Yat-sen University(中山大学)
  • Huawei Noah’s Ark Lab(华为诺亚方舟实验室)

机构由 AI 辅助整理,请以论文原文为准。

Hankz Hankui Zhuo, Shuting Deng, Mu Jin, Zhihao Ma, Kebing Jin, Chen Chen, Chao Yu

更新

AI总结:

本文提出一种结合符号选项与分层规划模型的深度强化学习框架,通过循环训练自动学习规划模型以提升数据效率、可解释性与可迁移性,并在两个领域验证了其有效性。

AI中文摘要:

尽管深度强化学习(DRL)在现实世界应用中取得了巨大成功,但它仍然面临三个关键问题,即数据效率低、缺乏可解释性和可迁移性。最近的研究表明,将符号知识嵌入DRL有望解决这些挑战。受此启发,我们提出了一种带有符号选项的新型深度强化学习框架。我们的框架具有循环训练过程的特点,该过程通过使用从交互轨迹中自动学习到的规划模型(包括动作模型和分层任务网络模型)和符号选项进行规划,从而指导策略的改进。学习到的符号选项减轻了对专家领域知识的密集需求,并为策略提供了内在的可解释性。此外,通过使用符号规划模型进行规划,可以进一步提高可迁移性和数据效率。为了验证我们框架的有效性,我们分别在Montezuma's Revenge和Office World两个领域进行了实验。结果表明,该框架具有可比的性能、更高的数据效率、可解释性和可迁移性。

英文摘要:

Despite of achieving great success in real-world applications, Deep Reinforcement Learning (DRL) is still suffering from three critical issues, i.e., data efficiency, lack of the interpretability and transferability. Recent research shows that embedding symbolic knowledge into DRL is promising in addressing those challenges. Inspired by this, we introduce a novel deep reinforcement learning framework with symbolic options. Our framework features a loop training procedure, which enables guiding the improvement of policy by planning with planning models (including action models and hierarchical task network models) and symbolic options learned from interactive trajectories automatically. The learned symbolic options alleviate the dense requirement of expert domain knowledge and provide inherent interpretability of policies. Moreover, the transferability and data efficiency can be further improved by planning with the symbolic planning models. To validate the effectiveness of our framework, we conduct experiments on two domains, Montezuma's Revenge and Office World, respectively. The results demonstrate the comparable performance, improved data efficiency, interpretability and transferability.

↑