arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

扩散策略改进:基于提案条件精炼流的方法

Diffusion Policy Improvement with Proposal-Conditioned Refinement Flows

Junhyun Ha, Juho Lee, Byoungwoo Park

arXiv 2609.36812首次发表:更新:

发表机构

KAIST(韩国科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出提案条件精炼流(PReFlow),结合评论家提案选择与条件精炼流,通过KL正则化目标优化策略提取,在50个OGBench任务上实现最优在线微调性能(91%)。

AI 中文摘要

扩散策略和流策略能够建模离线强化学习(RL)中的复杂行为。然而,对其与行为策略的KL散度进行惩罚可能会抑制那些具有高评论家值但行为密度低的动作。直接精炼行为提案可能是一种替代方案,但高斯或确定性编辑器限制了表达能力,无法为同一提案表示多个分离的模态。在这项工作中,我们引入了提案条件精炼流(PReFlow),这是一种策略提取方法,结合了基于评论家的提案选择与条件精炼流。为了同时优化提案选择和精炼,我们制定了一个KL正则化目标,其最优解在高斯平滑行为先验下诱导出一个关于最终动作的吉布斯策略。精炼流能够表示多个高值模态,而提案中心的高斯参考则调节大的动作变化。这种高斯参考进一步使我们能够利用无模拟的、闭式伴随匹配目标,从采样端点和评论家梯度中获取,从而产生一个单一的速度回归损失,无需反向伴随求解。在50个OGBench任务上,PReFlow在离线性能上具有竞争力,并在在线微调后获得所比较方法中最高的总分,在500K环境步后达到91%。

英文摘要

Diffusion and flow policies can model complex behaviors in offline reinforcement learning (RL). However, penalizing their KL divergence from the behavior policy can discourage actions having high critic values with low behavior density. Directly refining behavior proposals may be an alternative, yet Gaussian or deterministic editors limit expressiveness to represent multiple separated modes for the same proposal. In this work, we introduce Proposal-Conditioned Refinement Flows (PReFlow), a policy extraction method combining critic-based proposal selection with a conditional refinement flow. To optimize proposal selection and refinement together, we formulate a KL-regularized objective whose optimum induces a Gibbs policy over final actions under a Gaussian-smoothed behavior prior. The refinement flow can represent multiple high value modes, while a proposal-centered Gaussian reference regulates large action changes. This Gaussian reference further enables us to make use of simulation-free, closed form adjoint matching targets from sampled endpoints and critic gradients, yielding a single velocity regression loss without a backward adjoint solve. On 50 OGBench tasks, PReFlow achieves competitive offline performance and the highest aggregate score among the compared methods after online fine-tuning, reaching 91\% after 500K environment steps.

Comments27 pages, 10 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑