RL引导的PAC-NMPC用于未知环境中基于感知的概率安全导航
RL-Guided PAC-NMPC for Probabilistically-Safe Perception-Based Navigation in Unknown Environments
浏览论文内容
中文总结 AI 辅助
本文提出RL引导的PAC-NMPC方法,结合随机模型预测控制与强化学习,通过概率模型和硬约束实现未知环境中基于感知导航的概率安全,仿真和硬件实验验证了其安全性与扩展性。
中文摘要 AI 辅助
本文提出了一种结合随机非线性模型预测控制(SNMPC)与强化学习(RL)的方法,以实现未知环境中基于感知的概率安全导航。我们的方法首先利用RL训练概率性的演员-评论家模型和传感器预测模型。随后,我们将这些概率模型应用于基于采样的SNMPC框架中,该框架被称为“可能近似正确”(PAC)-NMPC,它使用硬约束来强制执行关于碰撞概率和值函数改进的有限时间统计保证。通过确保我们的有限时域SNMPC策略在期望上降低值函数,我们可以在满足概率安全约束的同时接近RL方法的长期性能。通过仿真实验,我们展示了我们的方法能够提高基于感知的RL导航策略的安全性,并扩展到具有大传感器输入空间和复杂非线性动力学的高维系统。我们还通过硬件实验展示了我们的方法,在未知环境中使用敏捷的固定翼飞行器进行基于视觉的导航时,性能得到了提升。
英文摘要
In this paper, we present an approach for combining stochastic nonlinear model predictive control (SNMPC) and reinforcement learning (RL) to enable probabilistically-safe perception-based navigation in unknown environments. Our method first uses RL to train probabilistic actor-critic and sensor prediction models. We then leverage these probabilistic models in a sampling-based SNMPC framework known as Probably Approximately Correct (PAC)-NMPC, which uses hard constraints to enforce finite-time statistical guarantees on the probability of collision and value function improvement. By ensuring that our finite-horizon SNMPC policies decrease the value function in expectation, we can approach the long-horizon performance of the RL approach while satisfying probabilistic safety constraints. Through simulation experiments, we show that our approach can improve the safety of perception-based RL navigation policies and scale to high dimensional systems with large sensor input spaces and complex nonlinear dynamics. We also demonstrate our approach through hardware experiments, showing improved performance for vision-based navigation with an agile fixed-wing aerial vehicle in unknown environments.