发表机构
Yonsei University; Seoul National University; University of Michigan; Korea Institute of Science and Technology (KIST)(延世大学; 首尔大学; 密歇根大学; 韩国科学技术研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出DODGER安全引导强化学习框架,直接执行策略动作并利用CBF过滤参考和约束违反塑造避碰行为,在仿真和真实人形实验中实现无运行时安全过滤器的动态障碍物导航。
AI 中文摘要
在以人为本的环境中运行的机器人必须在多个动态障碍物之间安全导航,以避免与人和周围基础设施发生碰撞。控制屏障函数(CBF)为安全过滤提供了一种有效机制,而近期基于CBF的强化学习(RL)方法将此类安全信息嵌入到学习策略中。然而,在训练期间仅执行安全过滤的动作会限制策略探索,这一限制在动态场景中尤为显著,因为安全性取决于机器人与障碍物的相对运动。我们提出DODGER,一种安全引导的强化学习框架,它直接执行策略生成的动作来驱动训练滚动,同时使用CBF过滤的参考和约束违反来塑造策略朝向避碰行为。我们通过Dubins车安全分析评估DODGER,并在全尺寸人形仿真和基于LiDAR感知的真实世界人形实验中,展示了在多个动态障碍物中的目标导向导航,无需运行时安全过滤器。
英文摘要
Robots operating in human-centered environments must safely navigate among multiple dynamic obstacles to avoid collisions with people and surrounding infrastructure. Control barrier functions (CBFs) provide an effective mechanism for safety filtering, and recent CBF-based reinforcement learning (RL) methods embed such safety information into learned policies. However, executing only safety-filtered actions during training can restrict policy exploration, a limitation that becomes particularly consequential in dynamic scenes where safety depends on relative robot-obstacle motion. We propose DODGER, a safety-guided RL framework that directly executes policy-generated actions to drive training rollouts while using CBF-filtered references and constraint violations to shape the policy toward collision-avoidance behavior. We evaluate DODGER through a Dubins-car safety analysis and demonstrate goal-directed navigation among multiple dynamic obstacles in full-order humanoid simulation and real-world humanoid experiments using LiDAR-based perception, without a runtime safety filter.
CommentsThe first three authors contributed equally to this work. Project Page: https://psh0823.github.io/dodger-homepage