arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06882cs.LGcs.AIcs.RO

离线强化学习中扩散策略的噪声空间策略梯度

Noisy-Space Policy Gradient for Diffusion Policies in Offline Reinforcement Learning

Mahmoud Selim, Cristina Cipriani, Karl H. Johansson

首次发表
浏览论文内容

中文总结 AI 辅助

针对扩散策略与强化学习整合的挑战,提出噪声空间Q函数和策略梯度(NSPG),通过KL正则化回归形式优化噪声潜在变量,在D4RL和OGBench基准上验证了有效性。

中文摘要 AI 辅助

扩散策略为连续控制提供了一种强大且富有表现力的参数化方式。然而,它们与强化学习的整合在概念上和算法上仍然具有挑战性。在这项工作中,我们通过引入一个噪声空间动作值(Q)函数来解决这一差距,该函数通过去噪过程引起的执行动作的分布为扩散潜在变量赋值。我们证明了这种构造具有精确的语义解释,并推导出噪声空间策略梯度(NSPG),该梯度仅使用干净的动作空间值估计来优化噪声潜在变量。基于这一结果,我们制定了在噪声潜在变量上的KL正则化策略改进,并表明所得目标具有扩散兼容的回归形式,避免了通过去噪过程进行反向传播。在基于状态的D4RL基准和基于视觉的OGBench任务上的实证结果表明,所提出的噪声空间目标为离线强化学习中训练扩散策略提供了原则性和有效的基础。项目网页:此https URL

英文摘要

Diffusion policies offer a powerful and expressive parameterization for continuous control. Yet, their integration with reinforcement learning remains conceptually and algorithmically challenging. In this work, we address this gap by introducing a noisy-space action-value (Q-)function that assigns values to diffusion latents through the distribution of executed actions induced by the denoising process. We show that this construction admits a precise semantic interpretation and derive a noisy-space policy gradient (NSPG) that optimizes noisy latents using only clean action-space value estimates. Building on this result, we formulate a KL-regularized policy improvement over noisy latents and show that the resulting objective admits a diffusion-compatible regression form, avoiding backpropagation through the denoising process. Empirical results on state-based D4RL benchmarks and vision-based OGBench tasks demonstrate that the proposed noisy-space objective provides a principled and effective basis for training diffusion policies in offline reinforcement learning. Project webpage: https://mahmoud-selim.github.io/NSPG/

发表机构

  • TRATON CV AB(特拉顿CV公司)
  • KTH Royal Institute of Technology(瑞典皇家理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑