LC-GRPO:通过朗之万校正弥合基于流的GRPO的训练-推理差距
LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction
浏览论文内容
中文总结 AI 辅助
LC-GRPO 是带朗之万校正的基于流的 GRPO 框架,通过对齐推理的 ODE 欧拉步加朗之万校正,缩小流模型训练与推理的样本差距,在多任务上提升奖励优化并保留生成质量。
中文摘要 AI 辅助
基于流的生成模型通常通过求解确定性常微分方程(ODE)进行采样,而在线强化学习需要随机 rollout 来进行策略探索和优化。因此,现有的用于流模型的 GRPO 方法在训练期间用随机微分方程(SDE)替代了推理时的 ODE。尽管 ODE 和 SDE 在连续时间内具有相同的边际分布,但它们的有限步离散化可能存在显著差异。特别是,随着探索噪声的增加,SDE rollout 通常会变得模糊,从而在强化学习使用的样本与测试时 ODE 采样器生成的样本之间造成不匹配。我们引入了 LC-GRPO,这是一种带有朗之万校正的基于流的 GRPO 框架。每个 rollout 转换首先执行与推理对齐的 ODE 欧拉步,然后应用随机朗之万校正,以针对所得时间步的边际分布。所需的评分直接从流速度中恢复,无需额外的评分模型,而所得转换仍为各向同性高斯分布,具有可处理的似然性,可用于策略优化。我们从理论上证明,在适当条件下,一次朗之万校正步骤可降低不完美 ODE 欧拉步的 Wasserstein 误差。在匹配的随机性水平下,我们进一步证明,所提出的转换比反向 SDE 的标准欧拉-丸山离散化更准确。在 SD3.5-Medium、FLUX.1-Dev 和 HunyuanVideo 上的实验表明,LC-GRPO 在文本到图像和文本到视频任务中持续改进了奖励优化,保留了生成质量,并大幅缩小了随机训练 rollout 与确定性测试时 ODE 推理之间的差距。
英文摘要
Flow-based generative models are typically sampled by solving a deterministic ordinary differential equation (ODE), whereas online reinforcement learning requires stochastic rollouts for policy exploration and optimization. Existing GRPO methods for flow models therefore replace the inference-time ODE with a stochastic differential equation (SDE) during training. Although the ODE and SDE share the same marginal distributions in continuous time, their finite-step discretizations can differ substantially. In particular, SDE rollouts often become blurry as the exploration noise increases, creating a mismatch between the samples used for reinforcement learning and those generated by the test-time ODE sampler. We introduce LC-GRPO, a flow-based GRPO framework with Langevin correction. Each rollout transition first takes an inference-aligned ODE Euler step and then applies a stochastic Langevin correction targeting the marginal distribution at the resulting timestep. The required score is recovered directly from the flow velocity, requiring no additional score model, while the resulting transition remains an isotropic Gaussian with a tractable likelihood for policy optimization. We theoretically show that, under suitable conditions, one Langevin correction step reduces the Wasserstein error of an imperfect ODE Euler step. At a matched randomness level, we further show that the proposed transition can be more accurate than the standard Euler--Maruyama discretization of the reverse SDE. Experiments on SD3.5-Medium, FLUX.1-Dev, and HunyuanVideo demonstrate that LC-GRPO consistently improves reward optimization across text-to-image and text-to-video tasks, preserves generation quality, and substantially narrows the gap between stochastic training rollouts and deterministic test-time ODE inference.