发表机构
Shanghai Artificial Intelligence Laboratory; Shanghai Innovation Institute; Fudan University(上海人工智能实验室; 上海创新研究院; 复旦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对扩散模型在线强化学习中的窗口选择、奖励饱和和样本效率问题,提出CAST方法,通过轨迹确定SDE窗口、因果场景图分解原子奖励及空间加权策略目标,在GenEval 2上显著超越Flow-GRPO。
AI 中文摘要
在线强化学习已被扩展到用于扩散模型(DM)图像生成的流匹配中。然而,这种范式面临三个局限性:(1)窗口选择。现有方法手动设置随机微分方程(SDE)采样窗口,即注入探索噪声的去噪步骤。我们则根据每个模型的去噪轨迹来确定该窗口。(2)奖励饱和。当前方法依赖于在人类标注上训练的评分模型;我们发现,在最新的SOTA开源DM上,这些分数极高且几乎无法区分,使得优势估计在很大程度上无效。(3)样本效率低下。单一标量奖励将不同的失败模式压缩为几乎相同的分数,留给针对性改进的梯度指导极少。为解决这些问题,我们提出了CAST(因果优势结构训练),一种针对预训练DM的RL微调方法,它(1)识别每个模型在图像中固定对象及其空间排列的去噪步骤,并利用该时机设置SDE窗口,(2)通过因果场景图(CSG)将每个提示分解为可验证原子,即最小的语义单元,如对象、数量、属性或空间关系,每个原子可独立检查,并分别奖励每个原子,以及(3)通过教师强制注意力将带符号的原子级优势投影到像素空间,并用它们对SDE策略目标进行空间加权。我们使用CAST微调了两个最强的开源DM,FLUX.2-dev和Qwen-Image-2512,并在组合基准GenEval 2上以及Qwen-Image-Bench上评估了它们的整体质量。在几乎相同的训练预算内,CAST在最具挑战性的GenEval 2提示上相对于基础模型的改进高达Flow-GRPO的3.07倍,同时整体生成质量也有所提升。
英文摘要
Online reinforcement learning has been extended to flow matching for diffusion model (DM) image generation. However, this paradigm faces three limitations: (1) Window selection. Existing methods manually set the stochastic differential equation (SDE) sampling window, i.e., the denoising steps where exploration noise is injected. We instead determine it from each model's denoising trajectory. (2) Reward saturation. Current methods rely on scoring models trained on human annotations; we find that such scores are extremely high and nearly indistinguishable on the latest SOTA open-source DMs, making advantage estimation largely ineffective. (3) Sample inefficiency. A single scalar reward collapses different failure modes into almost identical scores, leaving minimal gradient guidance for targeted improvement. To address these issues, we propose CAST (Causal Advantage-Structured Training), an RL fine-tuning method for pretrained DMs, which (1) identifies the denoising step at which each model fixes the objects and their spatial arrangement in the image and uses that timing to set the SDE window, (2) decomposes each prompt via Causal Scene Graphs (CSG) into verifiable-atoms, i.e., minimal semantic units such as an object, count, attribute, or spatial relation that can each be checked independently, and rewards each atom separately, and (3) projects the signed atom-level advantages into pixel space through teacher-forced attention and uses them to spatially weight the SDE policy objective. We fine-tune two of the strongest open-source DMs, FLUX.2-dev and Qwen-Image-2512, with CAST, and evaluate them on GenEval 2, a compositional benchmark, and on Qwen-Image-Bench for overall quality. Within almost the same training budget, CAST's improvement over the base model on the most challenging GenEval 2 prompts is up to 3.07x that of Flow-GRPO, while overall generation quality also improves.
CommentsProject page: https://opencausalab.github.io/CAST