arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.04282cs.CV

退一步以进千里:面向视觉生成的感知反思偏好优化

Step Back to Move Forward: Reflection-Aware Preference Optimization for Visual Generation

Junlong Wu, Jiuzhou Lin, Jia Sun, Boheng Zhang, Huaiqing Wang, Dewen Fan, Houde Liu, Qianqian Gan, Fan Yang, Tingting Gao

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对扩散模型偏好对齐中现有方法探索效率低、易陷入局部最优的问题,提出RA-GRPO框架,通过扩散反思与反事实路径合成提升性能,在T2I、T2V任务上表现优于现有方法。

中文摘要 AI 辅助

扩散模型已成为现代视觉生成的主流范式,大幅推动了多媒体内容合成的发展,尤其在文生图(T2I)和文生视频(T2V)任务中表现突出。为使这类生成模型更好地对齐人类偏好,强化学习(RL)近期作为一种后训练策略展现出强大潜力。然而,现有基于策略梯度的方法探索效率低下,易陷入局部最优,可能导致语义保真度和视觉真实感下降。为应对这些挑战,本文提出了感知反思组相对策略优化(Reflection-Aware GRPO, RA-GRPO),这是一种面向扩散生成模型的新型基于RL的偏好对齐框架。其核心思路是在优化过程中融入“反向”反思,以提升“正向”生成效果。我们首先引入扩散反思(Diffusion Reflection),通过使用弱估计器反转扩散过程来修正中间采样轨迹,引导隐状态向真实数据流形的高概率区域移动。此外,我们引入反事实路径合成(Counterfactual Path Synthesis),将这些修正后的轨迹隐式蒸馏到策略中,使模型能够内化基于搜索的探索的优势,同时不会产生推理时开销。在T2I和T2V模型上开展的大量实验表明,RA-GRPO的性能显著优于现有方法,尤其在缓解奖励黑客(reward hacking)和提升泛化能力方面表现突出。该方法仍保持架构无关性,可与标准流程无缝集成,为稳定的偏好对齐指明了有前景的方向。

英文摘要

Diffusion models have become the mainstream paradigm for modern visual generation and have substantially advanced multimedia content synthesis, especially in text-to-image and text-to-video tasks. To further align such generative models with human preferences, reinforcement learning (RL) has recently shown strong potential as a post-training strategy. Nevertheless, existing policy gradient-based methods often explore inefficiently, making them vulnerable to local optima that may degrade semantic faithfulness and visual realism. To address these challenges, we present Reflection-Aware GRPO (RA-GRPO), a new RL-based preference alignment framework for diffusion generative models. The core idea is to improve "forward" generation by incorporating "backward" reflection during optimization. We first introduce Diffusion Reflection, which rectifies intermediate sampling trajectories by inverting the diffusion process with a weak estimator, guiding latent states toward higher-probability regions of the true data manifold. Furthermore, we introduce Counterfactual Path Synthesis to implicitly distill these rectified trajectories into the policy, enabling the model to internalize the benefits of search-based exploration without incurring inference-time overhead. Extensive experiments on T2I and T2V models demonstrate that RA-GRPO significantly outperforms existing methods, particularly in mitigating reward hacking and improving generalization. The method remains architecture-agnostic and integrates seamlessly with standard pipelines, suggesting a promising direction for stable preference alignment.

发表机构

  • Tsinghua University(清华大学)
  • Kuaishou Technology(快手科技)

机构由 AI 辅助整理,请以论文原文为准。

↑