结合增量知识的神经符号推理用于样本高效分层强化学习
Neurosymbolic Reasoning with Incremental Knowledge for Sample Efficient Hierarchical Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
本研究提出结合增量知识(InK)的神经符号分层强化学习,通过符号高层组件执行符号规划、低层神经模块学习运动原语,在导航任务中显著提升了样本效率,还开发了信念世界树搜索方法。
中文摘要 AI 辅助
(平面)强化学习(RL)智能体在具有稀疏奖励且需要长程推理的环境中面临重大挑战。一种提高样本效率的有效方法是将知识融入学习与决策过程。在标准分层强化学习(HRL)中,知识以固定、不可更新的形式编码,例如架构选择,在整个学习过程中保持不变。采用固定HRL时,在获取足够环境知识前,无法利用探索过程中习得的增量知识进行推理,导致样本效率低下。本研究提出结合增量知识(Incremental Knowledge, InK)的神经符号HRL:符号高层组件在当前InK的可更新表示上执行符号规划(例如使用$D^*$),而低层目标条件神经模块通过奖励塑形学习运动原语。在导航任务上的实验表明,融入InK可显著提升样本效率。此外,为在给定世界先验知识下执行最优符号规划,我们开发了信念世界树搜索。代码可在该https网址获取。
英文摘要
(Flat) Reinforcement Learning (RL) agents face significant challenges in environments with sparse rewards that require long-horizon reasoning. A compelling approach to improve sample efficiency is to incorporate knowledge into learning and decision-making. In standard Hierarchical RL (HRL), knowledge is encoded in a fixed, non-updatable form, such as architectural choices, and remains unchanged throughout learning. With fixed HRL, reasoning with incremental knowledge learned during exploration is impractical before sufficient environmental knowledge is acquired, leading to poor sample efficiency. In this work, we propose neurosymbolic HRL with {\em Incremental Knowledge (InK)}: symbolic high-level components perform {\em symbolic planning} (e.g. using $D^*$) on an updatable representation of current InK, while low-level goal-conditioned neural modules learn motion primitives through experience using reward shaping. Experiments on navigation tasks demonstrate that incorporating InK substantially improves sample efficiency. Additionally, to perform {\em optimal} symbolic planning given {\em prior} knowledge about the world, we develop Belief World Tree Search. The code is available at https://github.com/CPS-research-group/ink_bwts.
发表机构
- NTU Singapore(新加坡南洋理工大学)
- IPAL, CNRS, France(法国IPAL-CNRS)
机构由 AI 辅助整理,请以论文原文为准。