发表机构
Aalborg University; Radboud University; Ruhr University Bochum(奥尔堡大学; 拉德堡德大学; 波鸿鲁尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对已知MDP转移图但转移概率未知的场景,提出将概率屏蔽与在线模型学习结合的自适应方法,通过实验验证其在安全强化学习中的有效性。
AI 中文摘要
概率屏蔽是一种用于安全强化学习(RL)的技术,通常,被称为“屏蔽”的静态观测器会将学习智能体的动作约束在那些能保证安全执行的动作范围内。传统上,屏蔽是基于底层马尔可夫决策过程(MDP)的转移概率计算得到的,因此,当无法预先给定MDP模型时,该技术便无法适用,而这恰恰是典型RL应用场景中的常见情况。在本文中,我们研究的问题是:在已知MDP的转移图但转移概率未知的情况下如何计算屏蔽。我们的方法将概率屏蔽与在线模型学习相结合:当RL智能体探索环境时,我们会估计转移概率,并基于该估计值计算屏蔽。尽管该屏蔽在初始阶段可能较为保守,但随着模型估计的精度不断提升,它会进行自适应调整,从而与RL智能体同步优化。这种自适应概率屏蔽的范式带来了诸多挑战,例如何时重新计算屏蔽,以及学习过程中如何在探索与安全之间取得平衡。我们在多个环境中对该范式的多种变体进行了实验评估。
英文摘要
Probabilistic shielding is a technique for safe reinforcement learning (RL). Typically, a static observer -- called the shield -- constrains the learning agent's actions to those for which acting safely remains feasible. Traditionally, the shield is computed from the transition probabilities of the underlying Markov decision process (MDP). Thus, this technique is not applicable when the MDP model is not given a priori, which, unfortunately, is the case in typical RL applications. In this paper, we study the problem of computing a shield in the setting where the transition graph of the MDP is known, but the transition probabilities are unknown. Our approach integrates probabilistic shielding with online model learning: as the RL agent explores the environment, we estimate the transition probabilities. From this estimate, we compute a shield. While the shield may be conservative initially, it adapts as the model estimate becomes more precise. Thus, the shield improves in tandem with the RL agent. This paradigm of adaptive probabilistic shielding raises a number of challenges, such as when to recompute the shield and how to balance between exploration and safety during learning. We empirically evaluate multiple variants of this paradigm across several environments.
Comments19 pages, 3 figures, 3 tables. To be published in the proceedings of RV 2026