具有时序逻辑规范的高效贝叶斯自适应强化学习
Efficient Bayes-Adaptive Reinforcement Learning with Temporal Logic Specifications
- University of Oxford(牛津大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出一种基于模型的贝叶斯自适应强化学习算法,通过同步LDBA与BAMDP并采用新颖的BAMCP规划,在未知环境中高效满足LTL规范,提升样本效率并减少训练违规。
AI中文摘要:
我们提出了一种新颖的端到端基于模型的强化学习(RL)算法,用于在未知环境中在给定线性时序逻辑(LTL)规范(例如安全性或可达性)下高效地合成策略。为此,LTL任务的极限确定性Büchi自动机(LDBA)表示与环境中的贝叶斯自适应马尔可夫决策过程(BAMDP)表示同步,这使我们能够利用通过贝叶斯RL实现的增强的探索-利用权衡,而非传统的非贝叶斯方法。我们进一步提出了一种新颖的贝叶斯自适应蒙特卡洛规划(BAMCP)算法,以在同步的BAMDP结构中实现近似贝叶斯最优策略合成。一系列有限和无限时域任务实验证明了我们的方法在属性满足和样本效率方面相对于传统无模型方法的有效性。额外的消融研究也成功突出了新颖的BAMCP算法相对于经典BAMCP在LTL任务满足方面的价值。最后,我们还展示了我们的方法在谨慎RL中的成功应用,即减少策略训练期间发生的任务违规次数。
英文摘要:
We present a novel end-to-end model-based Reinforcement Learning (RL) algorithm for efficient policy synthesis under given Linear Temporal Logic (LTL) specifications (e.g., safety or reachability) in unknown environments. To do so, a Limit-Deterministic B{ü}chi Automaton (LDBA) representation of the LTL task is synchronised with a Bayes-Adaptive Markov Decision Process (BAMDP) representation of the environment, which allows us to leverage an enhanced exploration-exploitation trade-off that is achieved via Bayesian RL, as opposed to traditional non-Bayesian approaches. We further propose a novel Bayes-Adaptive Monte-Carlo Planning (BAMCP) algorithm to allow for approximate Bayes-optimal strategy synthesis in the synchronised BAMDP construct. A range of finite- and infinite-horizon task experiments demonstrate the effectiveness of our approach in terms of both property satisfaction and sample efficiency, when compared to traditional model-free approaches. Additional ablation studies also successfully highlight the value of the novel BAMCP algorithm in comparison to classical BAMCP for LTL task satisfaction. Finally, we also showcase a successful application of our approach for \textit{cautious} RL, namely to reduce the number of task violations incurred during policy training.