发表机构
Arizona State University(亚利桑那州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对CBF短视和MPC计算昂贵的问题,提出BarrierFormer,一种障碍监督的Transformer框架,通过预测性滚动编码CBF约束学习无模型安全策略,实现实时安全导航,并在安全率和延迟上超越现有方法。
AI 中文摘要
控制障碍函数(CBF)已成为安全关键机器人领域中编码和强制状态约束最流行的工具之一。标准CBF方法本质上是短视的,因为它们仅在当前时间步强制安全性。因此,系统可能被驱动到安全集合的边界,而在未来时间步不存在可行的安全控制。基于模型预测控制(MPC)的方法通过在滚动时域上强制状态约束来解决这一问题。然而,此类方法通常需要已知模型以在每一步求解约束优化问题,这对于实时部署而言计算成本高昂。我们提出BarrierFormer,一种障碍函数监督的Transformer框架,通过在编码滚动级CBF约束中学习无模型安全策略来解决这些局限性。因果Transformer编码观测-动作历史,通过动力学头自回归生成预测性滚动以替代模型,并通过动作头向标称控制器提供残差校正以替代在线计算。作用于局部观测的障碍批评者评估沿该滚动的CBF约束违反情况,安全教师计算满足这些约束的障碍一致性动作,作为学习控制策略的直接监督目标。在推理期间,策略将观测-动作历史映射到控制动作,无需任何在线优化或模型知识,从而实现实时的无模型预测性安全强制。在用于安全目标导向导航的线性和非线性、2D和3D动力学系统上的评估表明,BarrierFormer在安全率和推理延迟方面优于现有的基于强化学习(RL)、基于扩散、基于MPC和基于Transformer的方法。
英文摘要
Control barrier functions (CBFs) have become one of the most popular tools for encoding and enforcing state constraints in safety-critical robotics. Standard CBF approaches are inherently myopic in nature as they enforce safety only at the current time step. Consequently, the system can be driven toward the boundary of the safe set where no feasible safe control exists at a future timestep. Model predictive control (MPC) based approaches address this by enforcing state constraints over a receding horizon. However, such approaches generally require the model to be known for solving a constrained optimization problem at every step, which is computationally expensive for real-time deployment. We propose BarrierFormer, a barrier-supervised transformer framework that addresses these limitations by encoding rollout-level CBF constraints in learning a model-free safe policy. A causal transformer encodes observation-action history, autoregressively generates a predictive rollout through the dynamics head to replace the model, and provides a residual correction to a nominal controller through the action head to replace the online computation. A barrier critic operating on local observations evaluates CBF constraint violations along this rollout, and a safety teacher computes barrier-consistent actions satisfying these constraints as direct supervision targets for the learned control policy. During inference, the policy maps observation-action history to control actions without any online optimization or model knowledge, enabling real-time model-free predictive safety enforcement. Evaluations across linear and nonlinear, 2D and 3D dynamical systems for safe goal-directed navigation demonstrate that BarrierFormer outperforms existing reinforcement learning (RL)-based, diffusion-based, MPC-based, and transformer-based approaches in safety rate and inference latency.
Comments23 pages, 2 figures, Accepted at 10th Conference on Robot Learning (CoRL 2026), Austin TX, USA