区域约束的群相对策略优化用于基于流的图像编辑
Region-Constrained Group Relative Policy Optimization for Flow-Based Image Editing
浏览论文内容
中文总结 AI 辅助
本文提出RC-GRPO-Editing框架,通过区域约束的GRPO后训练提升流基图像编辑中目标区域指令遵循和非目标区域保留。
中文摘要 AI 辅助
指令引导的图像编辑需要在目标修改与非目标保留之间取得平衡。最近,基于流的模型因其高保真度和高效的确定性ODE采样而成为指令引导图像编辑的强大基础。在此基础上,基于GRPO的奖励驱动后训练已被探索以直接优化编辑特定奖励,提高指令遵循和编辑一致性。然而,现有方法常面临噪声信用分配问题:全局探索也会扰动非目标区域,增加组内奖励方差并产生噪声GRPO优势。为了解决这个问题,我们提出了RC-GRPO-Editing,一种基于流的图像编辑下的区域约束GRPO后训练框架。它通过抑制背景诱导的噪声方差,实现更干净的局部信用分配,提高编辑区域指令遵循性同时保留非目标内容。具体而言,我们通过区域解耦的初始噪声扰动局部化探索,以减少背景诱导的奖励方差并稳定GRPO优势,并引入注意力集中奖励,使跨注意力与预期编辑区域在整个rollout过程中对齐,减少非目标区域的意外变化。在CompBench上的实验显示,在编辑区域指令遵循性和非目标保留方面取得了持续改进。
英文摘要
Instruction-guided image editing requires balancing target modification with non-target preservation. Recently, flow-based models have emerged as a strong and increasingly adopted backbone for instruction-guided image editing, thanks to their high fidelity and efficient deterministic ODE sampling. Building on this foundation, GRPO-based reward-driven post-training has been explored to directly optimize editing-specific rewards, improving instruction following and editing consistency. However, existing methods often suffer from noisy credit assignment: global exploration also perturbs non-target regions, inflating within-group reward variance and yielding noisy GRPO advantages. To address this, we propose RC-GRPO-Editing, a region-constrained GRPO post-training framework for flow-based image editing under deterministic ODE sampling. It suppresses background-induced nuisance variance to enable cleaner localized credit assignment, improving editing region instruction adherence while preserving non-target content. Concretely, we localize exploration via region-decoupled initial noise perturbations to reduce background-induced reward variance and stabilize GRPO advantages, and introduce an attention concentration reward that aligns cross-attention with the intended editing region throughout the rollout, reducing unintended changes in non-target regions. Experiments on CompBench show consistent improvements in editing region instruction adherence and non-target preservation.
发表机构
- South China Normal University(华南师范大学)
- South China Agricultural University(华南农业大学)
- Monash University(莫纳什大学)
机构由 AI 辅助整理,请以论文原文为准。