arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Flow-Map GRPO:基于锚定随机组合的少步流映射生成器的强化学习

Flow-Map GRPO: Reinforcement Learning for Few-Step Flow-Map Generators via Anchored Stochastic Composition

Zhiqi Li, Wen Zhang, Bo Zhu

arXiv 2607.00535首次发表:更新:

发表机构

Georgia Institute of Technology(佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出Flow-Map GRPO框架,通过锚定随机流映射组合(ASFMC)对确定性少步流映射生成器进行随机化,使其适用于强化学习后训练,并在FLUX文本到图像生成器上提升奖励和感知指标。

AI 中文摘要

少步流映射生成器,如一致性模型和MeanFlow,通过直接学习噪声与数据之间的长程传输映射来加速采样。然而,这些模型通常是确定性的,这使得它们难以使用需要随机轨迹和定义良好的似然比的强化学习(RL)后训练方法进行优化。现有的基于SDE的随机化技术是为具有无穷小或精细离散化过渡的基于速度的采样器设计的,因此不直接适用于长程流映射。在这项工作中,我们提出了Flow-Map GRPO,一个用于确定性少步流映射生成器的在线RL后训练框架。关键组件是锚定随机流映射组合(ASFMC),这是一种路径保持的随机化机制,通过基于锚定的条件重采样引入随机性,同时保留确定性流映射的原始边际概率路径。我们推导了单时间和双时间流映射参数化的GRPO目标。在基于FLUX的少步文本到图像生成器(包括MeanFlow和sCM)上的实验表明,Flow-Map GRPO在基于奖励、感知和任务级别的评估指标上改进了预训练的确定性流映射模型。我们的结果表明,确定性少步流映射生成器可以有效地与RL后训练对齐,而无需修改其原始模型参数化或将其重新训练为原生随机模型。

英文摘要

Few-step flow-map generators, such as consistency models and MeanFlow, accelerate sampling by learning long-range transport maps between noise and data. However, their deterministic transitions do not directly provide the stochastic trajectories and tractable likelihood ratios required by reinforcement learning (RL) post-training. Existing SDE-based stochasticization techniques target velocity-based samplers and do not directly extend to long-range flow-map transitions. We propose Flow-Map GRPO, an online RL post-training framework for deterministic few-step flow-map generators. Its key component, Anchored Stochastic Flow Map Composition (ASFMC), combines deterministic transport with anchor-based conditional resampling. We establish the conditions under which this construction preserves the marginal probability path and develop tractable local- and endpoint-anchor policies for two-time and single-time flow maps. These policies enable a unified GRPO training procedure. Experiments on FLUX-based MeanFlow and sCM generators demonstrate substantial improvements in OCR, PickScore, and GenEval at different numbers of inference steps, including joint OCR--PickScore gains with mixed rewards. Controlled ablations show that the stochastic transition design is essential for translating training rewards into generation quality. Flow-Map GRPO enables effective RL alignment of pretrained deterministic flow-map generators while retaining their original parameterization, without retraining them as native stochastic models.

Comments40 pages, 15 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑