arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

分数校准流:用于从非归一化密度采样及其在生成式在线强化学习中的应用

Score-Calibrated Flow for Sampling from Unnormalized Densities with Applications to Generative Online Reinforcement Learning

Zeyang Li, Yunan Wang, Risheek Garrepalli, Mohammad Ghavamzadeh, Navid Azizan

arXiv 2610.04696首次发表:更新:

发表机构

Massachusetts Institute of Technology; Qualcomm AI Research(麻省理工学院; 高通人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对在线强化学习中从非归一化密度采样训练生成式策略的难题,提出分数校准流(SCF)算法,利用自一致性和分数校准条件,无需重要性采样,在RL基准上匹配或超越现有方法并显著降低训练时间。

AI 中文摘要

扩散模型和流模型为在线强化学习(RL)提供了富有表现力的策略类别,能够实现多模态行为并提升性能。然而,训练这些策略仍然具有挑战性:评论家将期望策略指定为非归一化的玻尔兹曼密度,但并未提供直接来自该密度的样本。许多现有方法依赖重要性采样来构建训练信号,这可能导致高方差,增加计算成本并使训练不稳定。我们提出了分数校准流(SCF),一种简单且高效的算法,用于训练生成模型从非归一化密度中采样,无需重要性采样或通过采样轨迹进行反向传播。我们通过强制自一致性来学习期望的流,从而绕过了目标后验均值估计。通过联合利用给定的目标分数和流匹配的结构,我们将这些自一致性要求确立为分数校准的最优性条件,首先针对终端密度,然后针对可训练的 velocity 场。我们证明了它们的唯一解分别是目标密度和理想流模型,后者是条件流匹配(CFM)在目标样本可用时能够恢复的模型。我们将 velocity 条件表述为不动点方程,并利用其条件期望结构来构造一个停止梯度目标以强制执行该条件。由此产生的训练过程保留了 CFM 的可扩展的样本-插值-回归结构,尽管没有目标样本,而是使用当前流生成的端点。对于在线 RL,评论家梯度在生成的动作处提供目标分数,从而为 actor 训练提供了一种直接方法。在 RL 基准上的实验表明,SCF 匹配或优于最先进的生成式策略基线,同时大幅减少了训练时间。

英文摘要

Diffusion and flow models provide expressive policy classes for online reinforcement learning (RL), enabling multimodal behaviors and improved performance. However, training these policies remains challenging: the critic specifies the desired policy as an unnormalized Boltzmann density but does not provide direct samples from it. Many existing methods rely on importance sampling to construct training signals, which can suffer from high variance, increasing computational cost and destabilizing training. We propose Score-Calibrated Flow (SCF), a simple and efficient algorithm for training generative models to sample from unnormalized densities without importance sampling or backpropagation through the sampling trajectory. We learn the desired flow by enforcing self-consistency, bypassing target posterior mean estimation. By jointly exploiting the prescribed target score and the structure of flow matching, we establish these self-consistency requirements as score-calibrated optimality conditions, first for the terminal density and then for the trainable velocity field. We prove that their unique solutions are, respectively, the target density and the ideal flow model that conditional flow matching (CFM) would recover if target samples were available. We formulate the velocity condition as a fixed-point equation and exploit its conditional-expectation structure to construct a stop-gradient objective for enforcing it. The resulting training procedure retains the scalable sample-interpolate-regress structure of CFM despite the absence of target samples, using endpoints generated by the current flow. For online RL, the critic gradient supplies the target score at the generated actions, yielding a direct approach to actor training. Experiments on RL benchmarks demonstrate that SCF matches or improves upon state-of-the-art generative-policy baselines, while substantially reducing training time.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑