arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

安全分数匹配:基于Hamilton-Jacobi可达性的扩散策略用于在线安全强化学习

Safe Score Matching: Diffusion Policies with Hamilton-Jacobi Reachability for Online Safe Reinforcement Learning

Boyang Li, Matthew Kim, Sylvia Lee Herbert

arXiv 2609.33337首次发表:更新:

AI 中文总结

针对在线安全强化学习,提出安全分数匹配(SSM)方法,利用Hamilton-Jacobi可达性门控双分支分数目标,将Q分数匹配扩展到硬约束场景,在多个基准上实现低违规率与高性能。

AI 中文摘要

在线安全强化学习(RL)旨在最大化奖励的同时满足安全约束。安全强化学习中一个流行的研究方向是将安全约束放宽为软期望成本约束,并通过原始-对偶拉格朗日更新来解决由此产生的约束马尔可夫决策过程,但这种更新仅能保证平均意义上的安全。为解决这一局限,引入了硬性的、逐状态约束,并通常通过Hamilton-Jacobi(HJ)可达性来施加。然而,此类约束需要在可行区域和不可行区域中求解不同的目标:在可行区域中最大化奖励,在不可行区域中则恢复至可行区域。由此产生的目标动作分布本质上是多模态的,这种结构对现有基于HJ的安全强化学习中使用的高斯或确定性actor构成了根本性挑战,这些actor通常会坍缩到次优模态。扩散策略提供了表示此类分布所需的表达能力,而近期关于Q分数匹配的工作通过分数回归为在线强化学习训练扩散策略提供了一条途径——但仅应用于奖励最大化。我们提出了安全分数匹配(SSM),一种离策略actor-critic方法,通过使用HJ可达性门控双分支分数目标,将Q分数匹配适配到硬约束安全强化学习:在可行集内,去噪过程退化为对HJ critic判定为可行的动作进行Q分数匹配;在可行集外,恢复分支将去噪偏向于最坏情况违规较低的区域。在四旋翼和固定翼轨迹跟踪以及稳定-避障基准测试中,SSM以较低的假安全率取得了最佳或接近最佳的任务性能,而原始-对偶基线允许更多不安全行为,基于可达性的基线则往往更保守;在Safety-Gymnasium速度任务中,SSM以具有竞争力的奖励取得了最低成本。

英文摘要

Online safe reinforcement learning (RL) seeks policies that maximize reward while satisfying safety constraints. A popular line of research in safe RL relaxes safety to a soft expected-cost constraint and solves the resulting Constrained Markov Decision Process via primal-dual Lagrangian updates that only enforce safety on average. To address this limitation, hard, state-wise constraints are introduced and often imposed through Hamilton-Jacobi (HJ) reachability. Yet such constraints require solving different objectives in the feasible and infeasible regions: reward maximization in the former, recovery toward the feasible regions in the latter. The resulting target action distributions are inherently multimodal, and this structure poses a fundamental challenge for the Gaussian or deterministic actors used in existing HJ-based safe RL, which often collapse onto suboptimal modes. Diffusion policies provide the expressiveness needed to represent such distributions, and recent work on Q-score matching offers a route to training them for online RL by score regression -- but has been applied only to reward maximization. We propose Safe Score Matching (SSM), an off-policy actor-critic method that adapts Q-score matching to hard-constrained safe RL by gating a two-branch score target with HJ reachability: inside the feasible set, the denoising process degenerates to Q-score matching on actions classified as viable by the HJ critic; outside, a recovery branch biases denoising toward regions with lower worst-case violation. On quadrotor and fixed-wing trajectory-tracking and stabilize-and-avoid benchmarks, SSM attains the best or near-best task performance with low false-safe rates, whereas the primal-dual baseline admits more unsafe behavior and reachability-based baselines tend to be more conservative; on Safety-Gymnasium velocity tasks, SSM attains the lowest cost with competitive reward.

CommentsAccepted to NeurIPS 2026. 29 pages, 4 figures. Code: https://github.com/byli888/safe-score-matching

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑