arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.12784cs.ROcs.LG

安全强化学习中高效探索的方向约束

Directional Constraints for Efficient Exploration in Safe Reinforcement Learning

Paolo Magliano, Puze Liu, Jan Peters, Davide Tateo, Raffaello Camoriano

首次发表
浏览论文内容

中文总结 AI 辅助

针对安全强化学习中约束学习降低速度和性能的问题,提出扩展ATACOM框架的ATACOM-DC方法,通过引入方向约束区分动作,改善安全与性能权衡,在模拟机器人控制任务中评估,分析约束违反成本和任务性能。

中文摘要 AI 辅助

强化学习变革了机器人研究领域,能在模拟中强大地学习复杂机器人技能。但在开放环境中的实际部署需要强大安全保障。安全强化学习方法通过实施安全约束来满足此要求。然而,在约束下学习常降低学习速度并导致次优任务性能。为此,本文提出扩展ATACOM框架,其为可与现有强化学习算法集成以实施源自系统先验知识或直接从数据学习的约束的先进可靠安全层。所提方法ATACOM方向约束(ATACOM-DC)通过引入区分接近和远离约束边界的动作的方向约束显著改善了安全-性能权衡,仅在必要时激活约束实施。我们在模拟中的一系列具有挑战性的机器人控制任务上评估了该方法,分析了约束违反成本和实现的任务性能。

英文摘要

Reinforcement Learning has revolutionized the landscape of robotic research, allowing robust learning of complex robotic skills in simulation. However, real-world deployment in open-ended environments requires strong safety guarantees to prevent dangerous or harmful behaviors. Safe Reinforcement Learning methods address this requirement by enforcing safety constraints. Nevertheless, learning under constraints often reduces learning speed and could lead to suboptimal task performance, as the agent must solve a more complex constrained optimization problem compared to unconstrained settings. To tackle this issue, in this work, we propose an extension of the ATACOM framework, a state-of-the-art reliable safety layer that can be integrated with existing Reinforcement Learning algorithms to enforce constraints derived from prior knowledge of the system or learned directly from data. Our proposed method, named ATACOM Directional Constraints (ATACOM-DC), significantly improves the safety-performance trade-off by introducing directional constraints that distinguish between actions approaching and moving away from constraint boundaries, activating constraint enforcement only when necessary. We evaluate our method across a range of challenging robotic control tasks in simulation, analyzing both constraint-violation costs and achieved task performance. Code and additional material at https://atacom-dc.robot-learning.net.

发表机构

  • Dipartimento di Automatica e Informatica, Politecnico di Torino(自动控制与信息工程系,都灵理工大学)
  • Tongji University, Shanghai Research Institute for Intelligent Autonomous Systems(同济大学,上海智能自主系统研究院)
  • German Research Center for AI (DFKI)(德国人工智能研究中心(DFKI))
  • Intelligent Autonomous Systems Group, TU Darmstadt(智能自主系统组,达姆施塔特工业大学)
  • Lund University(隆德大学)
  • Istituto Italiano di Tecnologia(意大利技术研究院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑