强化学习中的计算高效安全探索
Computationally efficient safe exploration in reinforcement learning
浏览论文内容
中文总结 AI 辅助
本文提出一种基于Nadaraya-Watson估计器的计算高效安全探索算法CoLSafe-MDP,用于约束马尔可夫决策过程,相比高斯过程方法大幅降低计算成本,并在网格环境和火星地形数据上验证了性能。
中文摘要 AI 辅助
现实应用中的强化学习要求在探索过程中提供安全保障。典型的强化学习算法不提供此类保证,而许多提供保证的改进方法依赖于高斯过程(GPs),其计算成本很高。我们提出了一种基于Nadaraya-Watson估计器的计算轻量级算法,用于安全探索并优化约束马尔可夫决策过程(MDPs)。我们的算法\textsc{CoLSafe-MDP}使用一种估计器,其计算时间恒定且具有估计边界,相比基于高斯过程的同类方法(其计算复杂度随数据点数量呈立方增长)有显著改进。随后,我们在基于网格的环境和观测到的火星地形数据上评估了其性能。
英文摘要
Reinforcement learning in real-life applications requires safety guarantees during exploration. Typical reinforcement learning algorithms do not provide such guarantees, and many modifications that do rely on Gaussian processes (GPs), which have a large computational cost. We propose a computationally lightweight algorithm based on the Nadaraya-Watson estimator that safely explores and optimizes constrained Markov decision processes (MDPs). Our algorithm, \textsc{CoLSafe-MDP}, uses an estimator that scales in constant-time with bounds on the estimates, a significant improvement from its GP-based counterparts that scale cubically with the number of data points. We then evaluate its performance in a grid-based environment and on observational Martian terrain data.
发表机构
- Aalto University(阿尔托大学)
机构由 AI 辅助整理,请以论文原文为准。