基于降阶模型的类人机器人安全导航学习
Learning Safe Humanoid Navigation from Reduced Order Models
- California Institute of Technology(加州理工学院)
- Amazon Safe Autonomy Frontiers (SAF) Lab(亚马逊安全自主前沿实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对类人机器人多楼层导航难题,提出将导航分解为降阶模型导航与全阶动力学策略启动两阶段,并应用泊松安全过滤器,在Unitree G1上实现超10米垂直位移与100米路径的无地图导航。
AI中文摘要:
类人机器人研究在运动控制方面取得了快速进展,近期成果不断拓展自主导航的边界。我们证明,标准的单阶段强化学习导航流程难以扩展到多层及多楼层地形,其局限性源于类人机器人在楼梯等复杂地形交互中的困难。为克服这一挑战,我们将导航问题分解为两个部分。首先,我们训练一个基于降阶动力学但使用完整三维激光雷达观测的策略,以导航复杂的多楼层地形。随后,我们利用该导航知识来启动一个基于全阶类人动力学的策略,并在循环中冻结运动策略。此外,我们证明,对导航策略输出应用泊松安全过滤器,可在不降低导航成功率的情况下,恢复面对分布外障碍物时的安全性。我们在Unitree G1上展示了所得到的RoM-Nav策略,实现了无地图的多楼层导航,试验覆盖超过10米的垂直位移和超过100米的路径长度。项目页面附有视频,网址为https URL。
英文摘要:
Research in humanoid robotics has achieved rapid progress in locomotion, and recent results have pushed the boundary on autonomous navigation. We demonstrate that a standard single-stage RL navigation pipeline struggles to scale to multi-level and multi-story terrain, limited by the difficulty of complex humanoid terrain interactions such as stairs. To overcome this challenge, we decompose the navigation problem into two pieces. First, we train a policy operating on the reduced order dynamics but with full 3D LiDAR observations to navigate complex, multi-story terrain. We then utilize this navigation knowledge to kickstart a policy operating on the full-order humanoid dynamics, with a frozen locomotion policy in the loop. Additionally, we demonstrate that applying a Poisson safety filter to the navigation policy output recovers safety in the presence of out-of-distribution obstacles, without dropping navigation success rate. We demonstrate the resulting RoM-Nav policy on a Unitree G1, accomplishing mapless multi-floor navigation covering trials with over 10m of vertical displacement and over 100m of path length. Project page with videos https://wdc3iii.github.io/rom-nav/ .