arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.14087cs.ROcs.SYeess.SY

无梯度神经Hamilton-Jacobi可达性分析用于可扩展的安全关键控制

Gradient-Free Neural Hamilton-Jacobi Reachability for Scalable Safety-Critical Control

Zeyuan Feng, Ali Fuat Sahin, Santiago Thorup, Somil Bansal

首次发表
浏览论文内容

中文总结 AI 辅助

提出一种无梯度神经HJ可达性框架,利用bang-bang结构和窗口化时间课程学习值函数,在高达80维基准和16,000维F1-tenth赛车上实现可扩展的安全控制与零样本泛化。

中文摘要 AI 辅助

Hamilton-Jacobi (HJ) 可达性分析为安全关键机器人系统综合安全证书和鲁棒控制器提供了一个原则性框架。然而,将可达性分析应用于高维非线性系统仍然具有挑战性:经典的基于网格的求解器受制于维度灾难,连续时间神经求解器需要精确的空间值梯度,而基于强化学习的方法往往存在边界锚定薄弱和对抗性策略优化非平稳的问题。我们提出了一种针对控制-扰动仿射系统的离散时间神经可达性框架,通过Bellman-Isaacs值传播学习后向可达管(BRTs)和后向可达-避免管(BRATs)。我们的关键思想是将方程驱动的自监督与结构化策略学习相结合:不计算显式的偏微分方程梯度,而是利用最优安全干预的bang-bang结构,从无梯度值探测中构造近似教师动作,将对抗性actor学习转化为监督式策略学习。为了稳定长时域值传播,我们利用学到的actor,从终端边界开始,通过窗口化时间课程向后训练值函数,其中每个窗口用作下一个窗口的边界条件。在高达80维的基准问题上,我们的方法学习了准确的可达性值函数,同时相比现有的基于学习的求解器提高了稳定性。我们进一步在F1-tenth赛车上展示了超过16,000维自我中心输入的观测空间可扩展性。学习到的安全过滤器对未见过的赛道实现了零样本泛化,并迁移到实体遥控车上,实现了实时鲁棒避障。

英文摘要

Hamilton-Jacobi (HJ) reachability provides a principled framework for synthesizing safety certificates and robust controllers for safety-critical robotic systems. However, applying reachability analysis to high-dimensional nonlinear systems remains challenging: classical grid-based solvers suffer from the curse of dimensionality, continuous-time neural solvers require accurate spatial value gradients, and reinforcement-learning-based approaches often suffer from weak boundary anchoring and non-stationary adversarial policy optimization. We propose a discrete-time neural reachability framework for control-disturbance-affine systems that learns backward reachable tubes (BRTs) and backward reach-avoid tubes (BRATs) through Bellman-Isaacs value propagation. Our key idea is to combine equation-driven self-supervision with structured policy learning: rather than computing explicit PDE-gradients, we exploit the bang-bang structure of optimal safety interventions to construct approximate teacher actions from gradient-free value probes, converting adversarial actor learning into supervised policy learning. To stabilize long-horizon value propagation, we leverage the learned actor to train the value function backward from the terminal boundary using a windowed temporal curriculum, where each window is used as the boundary condition for the next window. Across benchmark problems up to 80 dimensions, our method learns accurate reachability value functions while improving stability over existing learning-based solvers. We further demonstrate observation-space scalability on F1-tenth racing with over 16,000-dimensional egocentric inputs. The learned safety filter generalizes zero-shot to unseen tracks and transfers to a physical RC car, achieving real-time robust collision avoidance.

发表机构

  • Stanford University(斯坦福大学)
  • École Polytechnique Fédérale de Lausanne(洛桑联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

↑