arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.24057cs.LG

使用后继表示的约束强化学习

Constrained Reinforcement Learning Using Successor Representations

Michael Girstl, Alexander Mattick, Christopher Mutschler

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对现实世界强化学习中策略难适应成本函数变化的问题,提出SafeDSR方法,通过引入可学习权重矩阵扩展深度后继表示到约束强化学习,能快速重训策略,在二维导航环境展示竞争力与灵活性。

中文摘要 AI 辅助

现实世界中的强化学习依赖于将安全约束纳入策略的能力。一种常见的建模此类约束的方法是在马尔可夫决策过程中引入额外的成本信号,该信号独立于奖励信号通知智能体不良行为。然而,当前方法难以适应例如领域转移或障碍物随时间移动所带来的成本函数变化。缺乏适应性意味着策略过于僵化,无法应对复杂的现实世界条件。我们提出了安全深度后继表示(SafeDSR),一种新方法,它允许针对新的成本结构快速重新训练策略。SafeDSR通过引入单个可学习权重矩阵将深度后继表示扩展到约束强化学习,以在动力学、奖励和成本之间解耦学习到的价值函数。如果环境的成本结构发生变化,该矩阵可以以监督方式更新,而不必调整整个网络。我们在可自由配置的二维导航环境中展示了这种能力,并表明我们的方法在简单导航任务中具有竞争力,同时更加灵活。

英文摘要

Real-world Reinforcement Learning depends on the ability to formulate safety constraints into a policy. A common way to model such constraints is to introduce an additional cost signal in the Markov Decision Process, which notifies the agent of unwanted behavior independently of the reward signal. Unfortunately, current methods are hard to adapt to changes in the cost function introduced by, e.g., domain shift or obstacles moving over time. The lack of adaptability means that policies are too unflexible to deal with complex real-world conditions. We propose the Safe Deep Successor Representation (SafeDSR), a novel method that allows quick retraining of policies towards new cost structures. SafeDSR extends the Deep Successor Representation (Kulkarni et al., 2016) to Constrained Reinforcement Learning by introducing a single learnable weight matrix to decouple the learned value function across dynamics, rewards, and costs. This matrix can be updated in a supervised manner instead of having to adapt the whole network if the cost structure of the environment changes. We demonstrate this ability in a freely configurable two-dimensional navigation environment and show that our method is competitive on a simple navigation task while being considerably more flexible

发表机构

  • Technical University of Darmstadt (TU Darmstadt)(达姆施塔特工业大学)
  • Hessian Center for Artificial Intelligence (hessian.AI)(黑森州人工智能中心)
  • Fraunhofer Institute for Integrated Circuits IIS, Fraunhofer IIS(弗劳恩霍夫集成电路研究所IIS)
  • University of Technology Nuremberg (UTN)(纽伦堡工业大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑