AI 中文总结
本文提出WTF框架,通过最优传输正则化将奖励微调转化为确定性最优控制问题,实现无模拟强化学习,在少步推理下提升奖励并减少280倍训练计算。
AI 中文摘要
奖励微调旨在更新预训练的基于流的生成模型,以提高其生成样本的下游奖励。现有方法通常将此问题表述为从奖励倾斜分布中采样,即KL正则化奖励最大化问题的解。在此,我们引入一种直接基于预训练漂移的最优传输正则化器。与KL奖励倾斜不同,所得目标将单个样本传输至更高奖励区域,而非重新加权基础分布。我们证明该问题等价于流上的确定性最优控制问题。给定预训练流映射,此等价性产生一种用于微调生成流的无模拟强化学习算法。我们将所得框架称为Wasserstein倾斜流映射(WTF),这是首个原生适用于流映射的端到端微调方案。输出为微调后的流映射,在少步推理预算下无需事后蒸馏即可保持强奖励对齐性能。在ImageNet-256和文本到图像上的实验表明,WTF在实现更高奖励的同时,多样性可与基线相当或更高,且训练计算量最多减少280倍。更广泛地,我们认为加速采样器(如流映射)是高效后训练的关键基础设施,且占主导地位的KL正则化公式仅是众多值得重新审视的选择之一。
英文摘要
Reward fine-tuning aims to update a pre-trained flow-based generative model to improve the downstream reward of its generated samples. Existing methods typically formulate this problem as sampling from a reward-tilted distribution, the solution to a KL-regularized reward-maximization problem. Here, we introduce an optimal transport regularizer built directly from the pre-trained drift. Unlike KL reward tilting, the resulting objective transports individual samples toward higher reward rather than reweighting the base distribution. We show that the resulting problem is equivalent to a deterministic optimal control problem on the flow. Given a pre-trained flow map, this equivalence yields a simulation-free reinforcement learning algorithm for fine-tuning generative flows. We call the resulting framework Wasserstein-Tilted Flow Maps (WTF), the first end-to-end fine-tuning recipe native to flow maps. The output is a fine-tuned flow map that retains strong reward-aligned performance at few-step inference budgets without post-hoc distillation. Experiments on ImageNet-256 and text-to-image show that WTF achieves higher reward with comparable or higher diversity than baselines, while requiring up to $280\times$ less training compute. More broadly, we argue that accelerated samplers such as flow maps are essential infrastructure for efficient post-training, and that the dominant KL-regularized formulation is only one of many choices worth revisiting.