arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多旋翼控制中安全强化学习的前向不变策略类

Forward-Invariant Policy Classes for Safe Reinforcement Learning in Multicopter Control

Chieh Tsai, Muhammad Junayed Hasan Zahed, Jinzhi Shen, Yi Xie, Ruoshan Lan, Majed Obaid, Salim Hariri, Hossein Rastgoftar

arXiv 2609.38655首次发表:更新:

发表机构

University of Arizona; Umm Al-Qura University(亚利桑那大学; 乌姆古鲁斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种基于前向不变策略类的强化学习安全增益调度框架,将安全性直接嵌入策略类,实现四旋翼悬停控制中无需运行时过滤的安全优化。

AI 中文摘要

本文提出了一种基于前向不变性诱导的动作空间设计的强化学习(RL)安全增益调度框架。在给定的标称模型和反演域假设下,该方法不通过运行时屏蔽或基于惩罚的约束来强制执行安全性,而是将安全性直接嵌入到策略类中。具体而言,我们构建了一个有限的反馈控制器库,这些控制器共享一个共同的李雅普诺夫证书,该证书在任意切换下建立指定可接受集合的前向不变性。因此,任何其动作被限制在该库中的策略,包括RL探索期间遇到的策略,都继承了相同的证书。因此,策略优化可以专注于闭环性能,而无需运行时安全过滤或动作投影。该框架在四旋翼悬停调节中实例化,其中DQN在认证的反馈控制器之间进行调度。非线性MuJoCo模拟展示了状态相关的增益调度,并实证评估了对风、模型失配、传感器噪声和传感延迟的鲁棒性。结果说明了如何通过在前向不变策略类上进行学习,将安全认证与策略优化分离。

英文摘要

This paper proposes a reinforcement learning (RL) framework for safe gain scheduling based on forwardinvariance-induced action-space design. Under stated nominalmodel and inversion-domain assumptions, rather than enforcing safety through runtime shielding or penalty-based constraints, safety is embedded directly into the policy class. Specifically, we construct a finite library of feedback controllers sharing a common Lyapunov certificate that establishes forward invariance of a prescribed admissible set under arbitrary switching. Consequently, any policy whose actions are restricted to this library, including policies encountered during RL exploration, inherits the same certificate. Policy optimization can therefore focus on closed-loop performance without runtime safety filtering or action projection. The framework is instantiated for quadcopter hover regulation, where a DQN schedules among certified feedback controllers. Nonlinear MuJoCo simulations demonstrate state-dependent gain scheduling and empirically evaluate robustness to wind, model mismatch, sensor noise, and sensing delay. The results illustrate how safety certification can be separated from policy optimization by learning over a forward-invariant policy class.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑