arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SGA-Flow-GRPO:用于Flow-GRPO的空间梯度引导信用分配

SGA-Flow-GRPO: Spatial Gradient-Guided Credit Assignment for Flow-GRPO

Yunkai Yang, Yudong Zhang, Xinying Chen, Bin Luo, Jienan Lyu, Kunquan Zhang, Weitao Wan, Runmin Dong

arXiv 2609.32340首次发表:更新:

AI 中文总结

针对Flow-GRPO在空间维度上统一分配优势导致次优策略的问题,提出基于奖励梯度的连续空间信用图及MAD鲁棒归一化方法,在GenEval上实现SOTA对齐质量。

AI 中文摘要

强化学习(RL)已被证明能有效将基于流的生成模型与人类偏好对齐。最近,Flow-GRPO作为一种高效的免评论家范式出现,通过对采样的候选轨迹计算优势值。然而,标准的Flow-GRPO在时间去噪步骤和空间潜在维度上应用统一的标量优势,没有明确考虑生成图像的空间结构,可能导致次优的策略更新。为解决这一问题,我们提出了一种针对扩散变换器(DiTs)的新型梯度引导空间信用分配框架。我们首先将Flow-GRPO中的转移级对数似然重新表述为与DiT补丁架构原生对齐的逐令牌表示,构建空间细粒度的重要性采样比率。为了在没有刚性、边界敏感的分割启发式的情况下分配局部信用,我们引入了一种由奖励梯度导出的连续空间信用图。关键的是,我们采用基于中位数绝对偏差(MAD)的异常值鲁棒归一化方案,并结合温度缩放,有效消除梯度噪声,同时突出功能性的提示对齐区域。在GenEval上的广泛评估表明,我们的方法实现了最先进的对齐质量,收敛速度与DiffusionNFT等顶级方法相当,同时在对齐性能上显著优于基于Flow-GRPO的方法。

英文摘要

Reinforcement Learning (RL) has proven effective in aligning flow-based generative models with human preferences. Recently, Flow-GRPO has emerged as an efficient critic-free paradigm by calculating advantages over sampled candidate trajectories. However, standard Flow-GRPO applies a uniform scalar advantage across both temporal denoising steps and spatial latent dimensions, without explicitly accounting for the spatial structure of generated images, which may lead to sub-optimal policy updates. To address this, we propose a novel gradient-guided spatial credit assignment framework tailored for Diffusion Transformers (DiTs). We first reformulate the transition-level log-likelihood in Flow-GRPO into a token-wise representation natively aligned with DiT patch architectures, constructing spatially fine-grained importance sampling ratios. To allocate localized credit without rigid, boundary-sensitive segmentation heuristics, we introduce a continuous spatial credit map derived from reward gradients. Crucially, we employ an outlier-robust normalization scheme based on Median Absolute Deviation (MAD) coupled with temperature scaling, effectively eliminating gradient noise while highlighting functional prompt-aligned regions. Extensive evaluations on GenEval show that our approach delivers SOTA alignment quality, achieving a convergence rate comparable to top-tier methods like DiffusionNFT while substantially improving upon Flow-GRPO-based methods in alignment performance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑