arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.04406cs.ROcs.SYeess.SY

可达性引导的序列二次规划保护的模型预测路径积分用于安全非线性预测控制

Reachability-Guided Sequential Quadratic Programming-Guarded Model Predictive Path Integral for Safe Nonlinear Predictive Control

Alexandre Didier, Jason J. Choi, Namhoon Cho, Claire J. Tomlin, Melanie N. Zeilinger

首次发表
浏览论文内容

中文总结 AI 辅助

提出ReSQ-MPPI,结合可达性分析与序列二次规划精炼,提升采样型模型预测控制在安全非线性环境中的安全性与性能。

中文摘要 AI 辅助

安全机器人控制通常需要将长期性能优化与硬状态和输入约束相结合,但现有方法往往仅部分解决这一权衡问题。基于采样的模型预测控制(MPC)方法,如模型预测路径积分(MPPI),在处理非线性和非凸环境方面有效,但其有限样本展开和无约束加权平均更新可能返回不安全的控制。确定性非线性MPC可以显式纳入约束,但其实时安全性和性能强烈依赖于热启动和局部收敛。Hamilton-Jacobi可达性(HJR)提供严格的安全证书,但离线值函数计算仅对降阶模型实用。我们提出ReSQ-MPPI,一种可达性信息驱动的生成-精炼架构,结合了这些互补优势。为降阶模型离线计算的HJR值函数引导在线MPPI采样朝向安全且有前景的轨迹候选。随后,MPPI解在全阶MPC问题中通过少量序列二次规划(SQP)迭代进行精炼。关键观察是,标准MPPI推断步骤是无约束加权最小二乘问题;ReSQ-MPPI将其替换为约束MPC精炼,在可行时恢复MPPI更新,否则最小化修改。在杂乱导航和自主赛车环境中的仿真表明,ReSQ-MPPI相比独立MPPI、MPC和可达性过滤的基于采样的控制基线,提高了安全性和性能。

英文摘要

Safe robot control often requires combining long-horizon performance optimization with hard state and input constraints, but existing approaches tend to address this tradeoff partially. Sampling-based model predictive control (MPC) methods such as model predictive path integral (MPPI) are effective in handling nonlinear and nonconvex environments, yet their finite-sample rollouts and unconstrained weighted-average update can return an unsafe control. Deterministic nonlinear MPC can explicitly incorporate constraints, but its real-time safety and performance depends strongly on warm starts and local convergence. Hamilton--Jacobi reachability (HJR) provides rigorous safety certificates, but offline value-function computation remains practical only for reduced-order models. We propose ReSQ-MPPI, a reachability-informed generation--refinement architecture that combines these complementary strengths. An HJR value function computed offline for a reduced-order model guides online MPPI sampling toward safe, promising trajectory candidates. The MPPI solution is then refined by a small number of sequential quadratic programming (SQP) iterations in a full-order MPC problem. A key observation is that the standard MPPI inference step is an unconstrained weighted least-squares problem; ReSQ-MPPI replaces it with a constrained MPC refinement that recovers the MPPI update when it is feasible and minimally modifies it otherwise. Simulations in cluttered navigation and autonomous racing environments demonstrate that ReSQ-MPPI improves safety and performance over standalone MPPI, MPC, and reachability-filtered sampling-based control baselines.

发表机构

  • ETH Zurich(苏黎世联邦理工学院)
  • University of California, Los Angeles(加州大学洛杉矶分校)
  • Seoul National University(首尔大学)
  • University of California, Berkeley(加州大学伯克利分校)

机构由 AI 辅助整理,请以论文原文为准。

↑