arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.14076eess.SYcs.SY

SafePG:具有硬约束的安全且全局最优的强化学习

SafePG: Safe and Globally Optimal Reinforcement Learning with Hard Constraints

Vipul K. Sharma, Wesley A. Suttle, S. Sivaranjani

首次发表
浏览论文内容

中文总结 AI 辅助

SafePG提出一种无模型策略梯度强化学习框架,通过构造随机包装策略并截断至硬安全约束,实现非线性系统安全控制,具备收敛性与全局最优性保证,并在四旋翼导航仿真中验证。

中文摘要 AI 辅助

我们提出了一种最优且收敛的无模型策略梯度(PG)强化学习(RL)框架,用于在硬安全约束下控制非线性动态系统。我们首先构造一类以确定性控制器为中心的随机包装策略,从而在未知环境中实现探索,同时保留底层确定性控制结构。然后,我们通过将这些随机策略截断到硬安全约束上,定义一类参数化的、通过构造保证安全的控制策略。接下来,我们通过测度论论证,证明了在截断策略类下可能非凸的RL目标及其策略梯度是良定义的。然后,我们开发了一种基于随机梯度上升的无模型PG算法,该算法直接搜索这些截断策略,并利用梯度优势建立收敛性和最优性保证。最后,我们通过在安全四旋翼导航问题上的仿真验证了该框架。

英文摘要

We present an optimal and convergent model-free policy gradient (PG) reinforcement learning (RL) framework for controlling nonlinear dynamical systems under hard safety constraints. We first construct a class of stochastic wrapper policies centered around a deterministic controller, thereby enabling exploration in unknown environments while preserving the underlying deterministic control structure. We then define a class of parameterized safe-by-construction control policies by truncating these stochastic policies onto hard safety constraints. We next establish, via measure-theoretic arguments, that the potentially nonconvex RL objective under the truncated policy class, as well as its policy gradients, are well-defined. We then develop a model-free PG algorithm based on stochastic gradient ascent that directly searches over these truncated policies and leverage gradient dominance to establish convergence and optimality guarantees. Finally, we validate this framework through simulations on a safe quadrotor navigation problem.

发表机构

  • Purdue University(普渡大学)
  • U.S. Army Research Laboratory(美国陆军研究实验室)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑