arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15988cs.RO

ResSafe:基于残差强化学习的人形机器人安全过滤

ResSafe: Learning Safety Filtering with Residual Reinforcement Learning for Humanoids

发表机构加州大学伯克利分校 · 加州大学洛杉矶分校
查看机构详情
  • University of California, Berkeley(加州大学伯克利分校)
  • University of California, Los Angeles(加州大学洛杉矶分校)

机构由 AI 辅助整理,请以论文原文为准。

Gechen Qu, Tong Zhang, Bike Zhang, Yen-Jen Wang, Koushil Sreenath, Claire Tomlin, Jason Jangho Choi

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出ResSafe,利用残差强化学习将性能与安全解耦,使残差策略作为隐式安全过滤器,改善人形机器人控制的性能-安全帕累托权衡。

中文摘要 AI 辅助

人形机器人的安全控制因其高维动力学、接触丰富的交互以及对扰动的敏感性而仍然具有挑战性。尽管强化学习已经实现了有效的运动控制和轨迹跟踪,但学习到的策略仍可能产生导致不稳定或跌倒的不安全动作。在这项工作中,我们提出将残差强化学习作为人形机器人安全控制的隐式安全过滤机制。我们不是依赖单一的标称策略来同时平衡性能、安全性和鲁棒性,而是将性能与安全性解耦。标称策略仅专注于任务性能,而残差策略学习安全修正。这种解耦带来了更好的性能-安全帕累托权衡,并避免了在单一策略训练中对多个相互竞争的奖励项进行仔细调优的需要。我们表明,残差策略可以充当隐式安全过滤器。

英文摘要

Safe control of humanoid robots remains challenging due to their high-dimensional dynamics, contact-rich interactions, and sensitivity to disturbances. Although reinforcement learning has enabled effective locomotion and motion tracking, learned policies can still generate unsafe actions that lead to instability or falls. In this work, we propose residual reinforcement learning as an implicit safety-filtering mechanism for safe humanoid control. Instead of relying on a single nominal policy to simultaneously balance performance, safety, and robustness, we decouple performance and safety. The nominal policy focuses solely on task performance, while a residual policy learns safety corrections. This decoupling leads to a better performance--safety Pareto trade-off and avoids the need for careful tuning of multiple competing reward terms within a single policy training. We show that the residual policy can act as an implicit safety filter.

↑