将安全转化为能力:通过安全过滤强化学习实现最小可被利用的机器人策略
Turning Safety into Competence: Minimally Exploitable Robot Policies via Safety-Filtered Reinforcement Learning
- Johns Hopkins University(约翰霍普金斯大学)
- Princeton University(普林斯顿大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对竞争性任务中机器人策略易被利用的问题,提出S2C两阶段框架,分离安全综合与任务学习,通过对抗性RL学习安全过滤器,在模拟和硬件测试中优于八个基线,实现高胜率与低可利用性。
AI中文摘要:
部署用于竞争性任务的机器人必须在超越对手的同时不牺牲安全性。现有方法,包括安全强化学习(RL),训练单一策略以同时实现任务成功和避免失败。这种耦合会使训练复杂化,并可能使学习到的策略易受蓄意攻击的利用。我们提出了“安全到能力”(S2C),一个两阶段的强化学习框架,将安全综合与竞争性任务学习分离。我们将竞争性交互建模为安全关键的马尔可夫博弈,并证明当所有参与者都致力于安全机动时,完美过滤能保持策略的不可利用性。S2C通过对抗性强化学习学习一个鲁棒的安全过滤器,在任务策略训练期间将其嵌入环境中,并在部署时保留相同的过滤器。在模拟触地得分游戏中,S2C优于八个安全强化学习基线,实现了最高的胜率和Elo评分,以及最低的可利用性。针对人类对手的硬件压力测试证实了S2C的能力。
英文摘要:
Robots deployed for competitive tasks must outmaneuver their opponents without sacrificing safety. Existing approaches, including safe reinforcement learning (RL), train a single policy to achieve task success and avoid failures simultaneously. This coupling can complicate training and leave the learned policy exploitable by deliberate attacks. We propose Safety to Competence (S2C), a two-stage RL framework that separates safety synthesis from competitive task learning. We formulate competitive interactions as safety-critical Markov games and prove that perfect filtering preserves policy non-exploitability when all players commit to safe maneuvers. S2C learns a robust safety filter via adversarial RL, embeds it in the environment during task policy training, and retains the same filter at deployment. In simulated touchdown games, S2C outperforms eight safe RL baselines, achieving the highest win rate and Elo rating, and the lowest exploitability. Hardware stress tests against a human opponent confirm S2C's competence.