面向安全强化学习的边界搜寻策略梯度算法
Boundary-Seeking Policy Gradient for Safe Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
本文提出BSPG算法,针对安全强化学习中标准梯度法易收敛到可行域内部的问题,结合切向与法向分量,在Safety-Gymnasium导航任务中实现更高奖励与更紧密的边界跟踪。
中文摘要 AI 辅助
安全强化学习是在满足安全约束的前提下最大化奖励。对于约束马尔可夫决策过程,基于占用度量的线性规划视角表明,当约束在最优状态下起作用时,最优策略恰好位于约束边界上,但标准的基于梯度的方法未利用该结构,常收敛到可行域内部。本文提出边界搜寻策略梯度(Boundary-Seeking Policy Gradient, BSPG),这是一种一阶方法,其更新结合了切向分量与符号残差驱动的法向分量:切向分量用于提升奖励同时将成本保持在一阶水平,法向分量则将策略从任一侧调节至起作用的边界;组合方向具有代数拉格朗日形式,含诱导系数且无需学习对偶变量。在精确梯度和给定正则条件下,约束残差从任一侧以有限时间范围的O(1/√T)界收敛至零,切向分量是边界上的奖励上升方向,任意收敛的参数序列在起作用的约束集上平稳,当极限也是可行集上的局部最大化器时满足KKT条件。这补充了现有分析,现有分析仅保证可行性却未刻画收敛时的约束值。在标准Safety-Gymnasium导航任务中,BSPG相比对比基准获得更高奖励,同时更紧密地跟踪边界。
英文摘要
Safe reinforcement learning maximizes reward subject to safety constraints. For Constrained Markov Decision Processes, the linear-programming view over occupancy measures implies that whenever the constraint is active at optimality, the optimal policy lies exactly on the constraint boundary, yet standard gradient-based methods do not exploit this structure and often settle in the feasible interior. We introduce Boundary-Seeking Policy Gradient (BSPG), a first-order method whose update combines a tangential component that improves reward while preserving cost to first order with a signed, residual-driven normal component that regulates the policy toward the active boundary from either side; the combined direction admits an algebraic Lagrangian form with an induced coefficient and no learned dual variable. Under exact gradients and stated regularity conditions, the constraint residual converges to zero from either side with a finite-horizon $O(1/\sqrt{T})$ bound, the tangential component is a reward-ascent direction on the boundary, and any convergent parameter sequence is stationary on the active constraint set, satisfying the KKT conditions when the limit is also a local maximizer over the feasible set. This complements existing analyses, which certify feasibility but do not characterize the constraint value at convergence. On a standard Safety-Gymnasium navigation task, BSPG attains higher reward while tracking the boundary more tightly than the compared baselines.
发表机构
- School of EECS, Washington State University(华盛顿州立大学电子工程与计算机科学学院)
机构由 AI 辅助整理,请以论文原文为准。