arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过基于加速度的CBF-QP约束执行实现强化学习策略在实际机器人部署中的安全执行

Safe Execution of RL Policies Via Acceleration-Based CBF-QP Constraint Enforcement for Real-World Robotic Deployments

Bastien Muraccioli, Alice Cariou, Pierre-Alexandre Leziart, Mathieu Celerier, Arnaud Demont, Gentiane Venture, Mehdi Benallegue

arXiv 2607.14488首次发表:更新:

发表机构

CNRS–AIST Joint Robotics Laboratory (JRL); AIST; The University of Tokyo(法国国家科学研究中心与日本产业技术综合研究所联合机器人实验室; 日本产业技术综合研究所; 东京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对强化学习缺乏安全保障限制硬件部署的问题,提出基于加速度的CBF-QP安全过滤器Acc-CBF-QP,能在运行时约束策略于安全集,减少约束违反,保留标称任务性能,在机器人实验中效果显著,且流程开源。

AI 中文摘要

强化学习能解决复杂机器人控制问题,但缺乏安全保障限制其硬件部署。特别是腿式机器人和操纵器常在安全关键边界附近操作,分布外状态会导致部署失败。为此,引入Acc-CBF-QP,一种基于加速度的二次规划安全过滤器,使用控制障碍函数在运行时将任何强化学习策略约束在安全集上,且不修改训练。该方法适用于无约束和安全强化学习策略,在统一优化框架内执行关节位置、速度、扭矩和碰撞约束。关键贡献是制定了RL+QP任务,在违反约束时调节与强化学习命令的偏差。引入了TorqueTask和Forward Dynamics Task,对安全性能权衡进行有原则的控制。在7自由度Kinova Gen3操纵器和19自由度Unitree H1人形机器人上进行实验,结果表明约束违反显著减少。在真实的H1硬件上,单独的安全强化学习策略每秒有10.04次违反,添加Acc-CBF-QP后减少92%至每秒0.80次。在Kinova Gen3上,Acc-CBF-QP完全消除了违反情况。在无违反情况下保留了强化学习目标的标称任务性能。在H1上的激进速度命令下,Acc-CBF-QP通过防止约束导致的关机提高了执行效果,延长了生存时间。整个流程是开源的。

英文摘要

Reinforcement Learning (RL) has demonstrated remarkable capabilities for solving complex robotic control problems, but its lack of safety guarantees severely limits deployment on hardware. In particular, as legged robots and manipulators often operate near safety-critical boundaries, out-of-distribution states can lead to failure upon deployment. To address this, we introduce Acc-CBF-QP, an acceleration-based Quadratic Program (QP) safety filter using Control Barrier Functions (CBFs) that constrains any RL policy onto a safe set at runtime without modifying training. The method applies to unconstrained and Safe-RL policies, and enforces joint position, velocity, torque, and collision constraints within a unified optimization framework. A key contribution is the formulation of RL+QP tasks that regulate deviation from the RL command when constraints would otherwise be violated. We introduce a TorqueTask, minimizing torque deviation, and a Forward Dynamics Task, minimizing induced acceleration deviation, thus providing principled control over safety-performance trade-offs. Experiments on a 7-DoF Kinova Gen3 manipulator and a 19-DoF Unitree H1 humanoid, both in simulation and on hardware, highlight substantial reductions in constraint violations. On the real H1 hardware, a Safe-RL policy alone yielded 10.04 violations/s, which were reduced by 92% to 0.80 violations/s when augmented with Acc-CBF-QP. On the Kinova Gen3, Acc-CBF-QP fully eliminated violations. Nominal task performance of the RL objective is preserved in violation-free regimes. Under aggressive velocity commands on H1, Acc-CBF-QP improves execution by preventing constraint-induced shutdowns, yielding longer survival times. The full pipeline is open-source.

Comments8 pages, 4 figures. Accepted for publication at IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑