发表机构
University of New South Wales (UNSW Sydney)(新南威尔士大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出AlignGraft方法,利用弱模型对的隐式奖励信号在测试时对齐冻结的大规模流模型,无需奖励函数,可跨规模迁移并提升偏好、组合和文本渲染性能。
AI 中文摘要
将文本到图像生成流模型与奖励对齐,使其遵循训练数据本身无法提供的目标。对齐微调通过强化学习(RL)或偏好优化实现这一目标,但必须对每个检查点重复进行,并且返回的模型固定在其训练时所使用的奖励和强度上。测试时对齐则在采样过程中引导冻结模型,允许任务特定和样本特定的引导。现有方法仅通过从奖励函数本身获取每步信号(通过其梯度或单独训练的价值函数)来实现这一点。我们提出改变监督来源:让一对弱模型而非奖励函数提供监督。一个源对齐模型与其基础模型一起作为源对齐对,将其训练奖励存储为采样器自身坐标中表达的隐式、逐步、KL锚定的信号。我们探索这种模型形式的监督能否跨越规模,并证明其可以:我们的方法AlignGraft通过在对齐过程中添加该对的速度差来对齐一个更大、冻结、从未微调的模型。在共享噪声核下,这种迁移是精确的,并且在测试时既不需要奖励也不需要其梯度。该方法没有调度,只有一个控制对齐强度的标量,并且可以将其外推到源对齐对的强度之外。在图像和视频流模型(Stable Diffusion 3.5、FLUX和Wan)上,这种迁移在偏好、组合和文本渲染奖励上提升了冻结的大模型,可以超过源对齐模型本身,并以较小的恒定采样开销保持大模型的保真度。大量实验表明,在弱模型上进行一次对齐运行产生的监督可以在测试时被整个模型家族重用。
英文摘要
Aligning a text-to-image generation flow model with a reward makes it follow objectives that the training data alone does not provide. Alignment fine-tuning delivers this by reinforcement learning (RL) or preference optimization, but it must be repeated for every checkpoint and returns a model fixed at the reward and strength it was trained with. Test-time alignment instead steers a frozen model during sampling, allowing task-specific and sample-specific guidance. Existing methods obtain this only by drawing the per-step signal from the reward function itself, through its gradient, or through a separately trained value function. We propose changing the supervision source: let a pair of weak models, not a reward function, supply the supervision. A source aligned model, kept together with its base as a source alignment pair, stores its training reward as an implicit, step-wise, KL-anchored signal expressed in the sampler's own coordinates. We explore whether this model-form supervision can cross scale, and show that it does: our method, AlignGraft, aligns a larger, frozen, never-tuned model by adding the pair's velocity difference during sampling. The transport is exact under a shared noising kernel and needs neither the reward nor its gradient at test time. The method has no schedules, only a single scalar that controls the alignment strength and can extrapolate it beyond that of the source alignment pair. Across image and video flow models (Stable Diffusion 3.5, FLUX, and Wan), the transfer lifts the frozen large model on preference, compositional, and text-rendering rewards, can exceed the source aligned model itself, and preserves the large model's fidelity at a small constant sampling overhead. Extensive experiments show that one alignment run on a weak model produces supervision that the whole model family can reuse at test time.