发表机构
University of Illinois Chicago; Texas A&M University(伊利诺伊大学芝加哥分校; 德克萨斯农工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究基于模型的安全强化学习,提出用控制障碍函数结合控制仿射随机傅里叶特征及自适应共形预测的框架,通过学习控制仿射动力学实现可验证安全策略,仿真结果验证了该框架在推车摆杆和3D四旋翼平台上的有效性。
AI 中文摘要
安全的基于模型的强化学习通常将控制理论分析与机器人的强化学习相结合,以安全地探索(部分)未知系统动力学,同时推导任务效率的控制动作。控制性能和安全保证通常依赖于部分建模的标称系统动力学的先验知识以及补偿残余模型不确定性的数据驱动模型。然而,现有方法往往忽略了残余模型不确定性的结构,这可能导致机器人行为过于保守或在基于安全学习的控制器下安全保证无效。本文提出了一种安全强化学习框架,该框架使用控制障碍函数(CBF)通过可验证的数据驱动安全策略学习控制仿射动力学。具体来说,我们首先使用控制仿射随机傅里叶特征(ARFF)以控制仿射形式对机器人动力学进行建模,这提供了与数据集大小成比例的计算效率,并减少了基于模型的强化学习的潜在模型偏差。然后,应用一种使用自适应共形预测(ACP)的无模型、高效不确定性量化方法来量化由学习到的控制仿射动力学引起的安全约束中的不确定性。这允许进行数据驱动的安全保证,适用于使用CBF进行有原则和高效的控制器合成。在推车摆杆和3D四旋翼平台上的仿真结果证明了所提出框架的有效性。
英文摘要
Safe model-based reinforcement learning (RL) often bridges control-theoretic analysis and RL for robots to safely explore (partially) unknown system dynamics while deriving control actions for task efficiency. The control performance and safety assurance typically rely on prior knowledge of partially modeled nominal system dynamics and the data-driven models that compensate for residual model uncertainties. However, existing methods often overlook the structure of residual model uncertainties (e.g., components affine in control), which could lead to overly conservative robot behaviors or invalid safety guarantees under the safe learning-based controllers. This paper proposes a safe reinforcement learning framework that learns control-affine dynamics with a certifiable data-driven safe policy using control barrier functions (CBF). Specifically, we first use Control-Affine Random Fourier Features (ARFF) to model robot dynamics in a control-affine form, which offers computational efficiency that scales with dataset size and reduces potential model bias for model-based reinforcement learning. Then, a model-free, efficient uncertainty quantification method using adaptive conformal prediction (ACP) is applied to quantify the uncertainty in the safety constraint arising from the learned control-affine dynamics. This allows for data-driven safety assurance amenable to principled and efficient controller synthesis with CBF. Simulation results on the cartpole and the 3D quadrotor platforms demonstrate the effectiveness of the proposed framework.
Comments8 pages, accepted to IROS 2026