arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

约束流策略更新:一种广义薛定谔桥视角

Constrained Flow Policy Updates: A Generalized Schrödinger Bridge View

Boyang Li, Matthew Kim, Sylvia Herbert

arXiv 2609.32952首次发表:更新:

AI 中文总结

针对在线安全强化学习中多模态动作分布与原始-对偶方法不稳定的问题,提出基于广义薛定谔桥的流策略更新方法RAFALE,直接微分增广目标并采用无密度动能正则化,在七个Safety-Gymnasium任务上实现奖励与安全约束的平衡。

AI 中文摘要

在线安全强化学习(RL)旨在寻找在满足安全约束的同时最大化奖励的策略。奖励与安全可能引发多模态动作分布,这对主流的原始-对偶方法构成挑战:高斯动作器可能坍缩到单一次优模态,且对非凸拉格朗日景观的优化可能不稳定。扩散策略和流策略能够表示此类分布,但近期采用扩散动作器的工作依赖于估计并匹配增广拉格朗日目标策略的得分。相反,我们直接通过流策略的生成路径对增广目标进行微分,因此无需估计得分。由于流策略缺乏现成的动作对数密度用于熵正则化,我们基于近期仅奖励方法FLAC的无密度动能正则化器,提出带最小能量的重参数化增广拉格朗日流动作器(RAFALE),一种用于安全RL的离策略动作器-评论家方法。我们将其更新表述为受约束的单端广义薛定谔桥,并证明对于每次源采样,该路径空间问题恰是动作空间中的熵正则化问题。在正噪声下,其解仅在估计成本超过由拉格朗日乘子设定的阈值处重新加权仅奖励动作分布。当噪声消失时,最优值收敛到流策略直接优化的最小能量映射目标的值。在七个Safety-Gymnasium任务中,RAFALE在每个任务上实现了具有竞争力的奖励,且平均最终成本在预算内,而强基线则在两者之间权衡;消融实验支持其增广目标和流动作器的必要性。

英文摘要

Online safe reinforcement learning (RL) seeks policies that maximize reward while satisfying safety constraints. Reward and safety can induce multimodal action distributions, challenging the prevailing primal-dual methods: Gaussian actors may collapse onto a single suboptimal mode, and optimization over the nonconvex Lagrangian landscape can be unstable. Diffusion and flow policies can represent such distributions, but recent work with a diffusion actor relies on estimating and matching the score of an augmented-Lagrangian target policy. Instead, we differentiate the augmented objective directly through the generation path of a flow policy, so no score needs to be estimated. Because a flow policy lacks a readily available action log-density for entropy regularization, we build on the density-free kinetic-energy regularizer of FLAC, a recent reward-only method, and propose Reparameterized Augmented-Lagrangian Flow Actor with Least Energy (RAFALE), an off-policy actor-critic method for safe RL. We formulate its update as a constrained one-ended generalized Schrödinger bridge and show that, for each source draw, this path-space problem is exactly an entropy-regularized problem in action space. At positive noise, its solution reweights the reward-only action distribution only where the estimated cost exceeds a threshold set by the Lagrange multiplier. As the noise vanishes, the optimal value converges to that of a least-energy map objective that the flow policy optimizes directly. Across seven Safety-Gymnasium tasks, RAFALE achieves competitive reward with mean final cost within budget on every task, whereas strong baselines trade one for the other; ablations support the necessity of both its augmented objective and its flow actor.

Comments24 pages, 5 figures, 11 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑