发表机构
Robotics and Mechanisms Laboratory (RoMeLa), Department of Mechanical and Aerospace Engineering, University of California, Los Angeles(机器人与机构实验室(RoMeLa),机械与航天工程系,加利福尼亚大学洛杉矶分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对人形机器人在动态环境中的导航问题,提出RAVEN框架,通过强化学习自适应调整可见性图规划器几何结构,结合无碰撞MPC层跟踪轨迹,经实验评估,该方法能减少超调、提高狭窄通道稳健性及在延迟噪声下可靠导航。
AI 中文摘要
动态环境中的人形机器人导航需要长视野规划并遵守短视野动态和安全约束。经典可见性图规划器与模型预测控制(MPC)结合可有效生成无碰撞轨迹,但其性能依赖手动调参和精确系统建模。实际机器人系统中,控制延迟、状态估计噪声和运动不确定性会导致超调及违反约束。我们提出RAVEN,一种用于稳健人形机器人导航的分层强化学习(RL)-MPC框架。与以往方法不同,RAVEN利用RL通过修改障碍物膨胀和相关图参数来调整可见性图规划器的几何结构。通过直接重塑自由空间几何,学习到的规划器改变全局路径拓扑以补偿延迟和跟踪缺陷。无碰撞MPC层随后跟踪规划轨迹并明确执行速度边界和避障约束。通过在现实延迟和观测噪声下训练,RAVEN学习到能提高稳健性的规划调整,同时保留显式长视野几何规划和约束优化,与端到端学习方法不同。我们将RAVEN与手动调参的可见性图MPC基线和纯RL导航策略进行评估。结果表明在障碍物附近超调减少,在狭窄通道中稳健性提高,在延迟和噪声下导航更可靠。这些发现表明强化自适应图构建与约束MPC相结合为稳健人形机器人导航提供了一种有效且可解释的替代端到端学习的方法。
英文摘要
Humanoid navigation in dynamic environments requires long-horizon planning while respecting short-horizon dynamic and safety constraints. Classical visibility-graph planners combined with model predictive control (MPC) can efficiently generate collision-free trajectories, but their performance depends on manually tuned parameters and accurate system modeling. In real robotic systems, control delays, state-estimation noise, and locomotion uncertainties can cause overshoot and constraint violations even when the nominal path is geometrically optimal. We propose RAVEN, a hierarchical reinforcement learning (RL)-MPC framework for robust humanoid navigation. Unlike prior approaches that use learning to tune cost weights or replace planning entirely, RAVEN employs RL to adapt the geometric construction of a visibility-graph planner by modifying obstacle inflation and related graph parameters. By directly reshaping the free-space geometry, the learned planner alters the topology of the global path to compensate for delay and tracking imperfections. A collision-free MPC layer then tracks the planned trajectory while explicitly enforcing velocity bounds and obstacle-avoidance constraints. By training under realistic delays and observation noise, RAVEN learns planning adaptations that improve robustness while retaining explicit long-horizon geometric planning and constrained optimization, in contrast to end-to-end learning approaches. We evaluate RAVEN against a manually tuned visibility-graph MPC baseline and a pure RL navigation policy. Results demonstrate reduced overshoot near obstacles, improved robustness in narrow passages, and more reliable navigation under delay and noise. These findings indicate that reinforcement-adaptive graph construction combined with constrained MPC provides an effective and interpretable alternative to end-to-end learning for robust humanoid navigation.