arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33444cs.LGcs.CV

阐明基于回归的扩散强化学习的设计空间

Elucidating the Design Space of Regression-based Diffusion Reinforcement Learning

Toyota Li, David Zhao, Alan Zhao

首次发表
浏览论文内容

中文总结 AI 辅助

本文统一了基于回归的扩散强化学习方法,提出新范式DiffusionRFT,收敛更快、训练更稳定、性能最优。

中文摘要 AI 辅助

一类新兴的方法族,放弃策略梯度,转而重新加权监督回归,在扩散和流模型的强化学习中获得了动力。DiffusionNFT、FlowAWR和RAM是具有对比动机的代表性机制。它们之间共享什么(如果有的话)尚不明确。我们证实,每个方法都是一个散度约束的奖励最大化问题的解,它们仅由定义约束的凸生成器区分。在统一建模框架下,我们揭示了先前工作在构建优势嵌入回归目标时所做的松弛:对于线性和指数倾斜形状,分别近似KKT条件和后验归一化器,而保留对概率单纯形的精确sparsemax投影用于线性倾斜,在本工作中产生了另一种优越的模型类型。在理论基础之外,我们进一步实证研究了设计空间,并阐明了回归式扩散RL的训练配方。保留我们在探索中发现的优点,产生了我们的范式DiffusionRFT,它收敛更快,训练更稳定,并达到最佳性能。

英文摘要

A nascent family of methods that forgoes the policy gradient and reweights a supervised regression instead has garnered momentum in reinforcement learning for diffusion and flow models. DiffusionNFT, FlowAWR, and RAM are representative regimes with contrasting motivations. It is yet opaque what, if anything, they share. We substantiate that each is the solution of one divergence-constrained reward-maximization problem, and they are differentiated only by the convex generator that defines the constraint. Under the unified modeling framework, we unravel the relaxations that prior art made during building the advantage-embedded regression target: approximating the KKT condition and posterior normalizer for the linear and exponential tilt shapes DiffusionNFT and FlowAWR respectively, while preserving the exact sparsemax projection onto the probability simplex for linear tilt leads to another superior model type in this work. Beyond the theoretical underpinnings, we further empirically investigate the design space and shed light on the training recipe for regression-style diffusion RL. Retaining the merits discovered during our exploration gives rise to DiffusionRFT, our paradigm that converges faster, trains more stably, and attains the top performance.

发表机构

  • Tencent(腾讯)

机构由 AI 辅助整理,请以论文原文为准。

↑