发表机构
University of Washington(华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对逆强化学习不安全及控制障碍函数设计难的问题,通过将奖励函数候选限制在CBF空间,实现安全在线控制与经验改进,能从无标签观察中恢复障碍函数,模拟实验显示其安全性能提升,并研究了不同IRL方法的权衡。
AI 中文摘要
逆强化学习(IRL)算法是从专家示范中学习和泛化的强大工具,但通常依赖无约束探索,对实际部署不安全。同时,控制障碍函数(CBF)可保证控制系统安全,但其解析设计耗时且深奥。本文通过在IRL中将奖励函数候选限制在CBF空间来共同解决这些限制,实现具有持续经验改进的安全在线控制。关键是,该框架能直接从无标签专家观察中数据驱动恢复障碍函数。实验表明,恢复的障碍函数对专家数据中完全不存在的不安全状态具有鲁棒性,在模拟导航环境中安全性能优于标准IRL基线,并研究了基于规划与基于策略的IRL方法在模拟和现实世界避障任务中的权衡。
英文摘要
Inverse Reinforcement Learning (IRL) algorithms are powerful tools for learning from and generalizing expert demonstrations, but they often rely on unconstrained exploration, rendering them unsafe for real-world deployment. Meanwhile, Control Barrier Functions (CBFs) can guarantee the safety of control systems, but the analytical design of CBFs can be time-consuming and esoteric. In this work, we address these limitations jointly by constraining reward function candidacy during IRL to the space of CBFs, yielding a formulation that exhibits safe online control with continuous experiential improvement. Crucially, this framework enables the data-driven recovery of barrier functions directly from unlabeled expert observations. We demonstrate that the recovered barrier function is robust to unsafe states entirely absent from the expert data. Furthermore, we benchmark our method against standard IRL baselines in a simulated navigation environment, demonstrating improved safety performance. Finally, we investigate the trade-offs of planning-based versus policy-based IRL methods across both simulation and a real world obstacle avoidance task.
Comments20 pages, 5 figures