arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19449cs.ROcs.SYeess.SY

赢得必胜之局:面向高维黑盒系统的严格可达-避免-驻留控制屏障函数

Winning a Won Game: Strict Reach-Avoid-Stay Control Barrier Functions for High-Dimensional Black-Box Systems

Donggeon David Oh, Duy P. Nguyen, Gongkai Yuan, Qingchen Li, Jaime Fernández Fisac, Haimin Hu

首次发表
浏览论文内容

中文总结 AI 辅助

针对高维黑盒系统,提出结合驻留与可达-避免价值的严格可达-避免-驻留Q控制屏障函数安全过滤器,在无需模型下保证安全到达并驻留,经四足跳跃和F1TENTH实验验证。

中文摘要 AI 辅助

机器人必须完成任务并维持已取得的成果,同时始终避免安全故障。严格可达-避免-驻留(sRAS)形式化了这一要求:安全地到达目标,并在首次进入后无限期地驻留其中。我们提出了一种针对有界不确定性下高维黑盒系统的sRAS Q控制屏障函数(CBF)安全过滤器。我们的构造将编码目标子集中安全永久驻留的驻留价值与编码该子集安全可达性同时避免无法保证安全永久驻留的目标状态的可达-避免价值相结合。我们证明了这些价值共同产生一个有效的鲁棒离散时间CBF,并将其提升为用于运行时干预的状态-动作Q函数。对于精确价值且在测度为零的条件下,我们的过滤器从几乎所有可赢的初始状态保持sRAS可行性,并在首次进入后将系统安全地保持在目标内,抵御所有允许的不确定性实现。我们采用基于可达性的对抗强化学习,仅利用黑盒交互即可实现可扩展的价值近似。值得注意的是,我们的过滤器的综合和部署都不需要已知动力学、仿射结构、价值导数或手工设计的屏障。我们在仿真和硬件中的四足机器人跨越间隙跳跃中验证了我们的框架,机器人跨越间隙、安全着陆并在此后保持安全。模拟的F1TENTH比赛进一步展示了安全超车和保持领先。

英文摘要

Robots must complete their tasks and maintain the achieved outcomes while avoiding safety failures at all times. Strict reach-avoid-stay (sRAS) formalizes this requirement: safely reaching a target and remaining there indefinitely after first entry. We propose an sRAS Q-control barrier function (CBF) safety filter for high-dimensional black-box systems under bounded uncertainty. Our construction combines a stay value encoding safe permanent residence in a target subset with a reach-avoid value encoding safe reachability of this subset while avoiding target states from which safe permanent residence cannot be guaranteed. We prove that these values jointly yield a valid robust discrete-time CBF and lift them to state-action Q-functions for runtime intervention. For exact values and under a measure-zero condition, our filter preserves sRAS feasibility from almost every winnable initial state and keeps the system safely within the target after first entry, against all admissible uncertainty realizations. We adopt reachability-based adversarial reinforcement learning for scalable value approximation using only black-box interactions. Notably, neither synthesis nor deployment of our filter requires known dynamics, affine structure, value derivatives, or hand-designed barriers. We validate our framework in quadruped gap jumping in simulation and hardware, where the robot crosses the gap, lands safely, and remains safe afterward. Simulated F1TENTH races further demonstrate safe overtaking and lead retention.

发表机构

  • Princeton University(普林斯顿大学)
  • Johns Hopkins University(约翰霍普金斯大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑