arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22919cs.LGcs.ROcs.SYeess.SY

强化学习中的计算高效安全探索

Computationally efficient safe exploration in reinforcement learning

Shreeram Murali, Shankar A. Deka, Dominik Baumann

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出一种基于Nadaraya-Watson估计器的计算高效安全探索算法CoLSafe-MDP,用于约束马尔可夫决策过程,相比高斯过程方法大幅降低计算成本,并在网格环境和火星地形数据上验证了性能。

中文摘要 AI 辅助

现实应用中的强化学习要求在探索过程中提供安全保障。典型的强化学习算法不提供此类保证,而许多提供保证的改进方法依赖于高斯过程(GPs),其计算成本很高。我们提出了一种基于Nadaraya-Watson估计器的计算轻量级算法,用于安全探索并优化约束马尔可夫决策过程(MDPs)。我们的算法\textsc{CoLSafe-MDP}使用一种估计器,其计算时间恒定且具有估计边界,相比基于高斯过程的同类方法(其计算复杂度随数据点数量呈立方增长)有显著改进。随后,我们在基于网格的环境和观测到的火星地形数据上评估了其性能。

英文摘要

Reinforcement learning in real-life applications requires safety guarantees during exploration. Typical reinforcement learning algorithms do not provide such guarantees, and many modifications that do rely on Gaussian processes (GPs), which have a large computational cost. We propose a computationally lightweight algorithm based on the Nadaraya-Watson estimator that safely explores and optimizes constrained Markov decision processes (MDPs). Our algorithm, \textsc{CoLSafe-MDP}, uses an estimator that scales in constant-time with bounds on the estimates, a significant improvement from its GP-based counterparts that scale cubically with the number of data points. We then evaluate its performance in a grid-based environment and on observational Martian terrain data.

发表机构

  • Aalto University(阿尔托大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑