arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

我们能否在无编辑奖励的情况下进行图像编辑的在线强化学习?

Can We Perform Online RL for Image Editing without Editing Rewards?

Qichao Ma, Jikang Cheng, Ling Liang, Zhaofei Yu, Tiejun Huang, Renye Yan

arXiv 2608.22780首次发表:更新:

AI 中文总结

本文提出Lever-Edit框架,将文本到图像(T2I)奖励生态系统迁移至图像编辑,在无编辑奖励的情况下优化编辑策略,实验效果优于直观迁移基线且接近基于编辑奖励的微调。

AI 中文摘要

强化学习(RL)可通过编辑特定奖励实现图像编辑的直接偏好优化,但这类奖励因三重监督成本高、任务相关校准复杂而发展不足。相比之下,文本到图像(T2I)生成受益于成熟多样的奖励生态系统,涵盖语义对齐、美学、真实感、字形及其他视觉偏好。将该生态系统扩展至图像编辑,可大幅拓宽基于RL的优化可访问的视觉偏好范围,由此引出核心问题:我们能否在无编辑奖励的情况下进行图像编辑的RL?本文提出,标准图像编辑维度有潜力映射至T2I奖励空间:图像质量可直接迁移,提示跟随可通过期望视觉状态的描述对齐,参考一致性可通过编码源内容以保留来实现粗略语义转换。然而,编辑指令指定相对变化,而T2I奖励需自包含的目标描述;此外,通用视觉-语言模型生成的语义有效描述可能与冻结奖励不兼容。因此,我们进一步引入Lever-Edit,这一两阶段框架学习用于反事实目标描述的奖励对齐描述器,冻结该描述器后,仅用迁移的T2I奖励优化编辑策略。实验表明,与基于编辑奖励的微调相比,该方法在编辑对齐和源保留方面具有竞争力,且优于直观迁移基线。

英文摘要

Reinforcement learning (RL) enables direct preference optimization for image editing through editing-specific rewards, which remain less developed due to costly triplet supervision and complex task-dependent calibration. In contrast, text-to-image (T2I) generation benefits from a mature and diverse reward ecosystem spanning semantic alignment, aesthetics, realism, glyph shape, and other visual preferences. Extending this ecosystem to image editing would substantially broaden the range of visual preferences accessible to RL-based optimization, prompting the central question: \emph{Can We Perform Image Editing RL without Editing Rewards?} In this paper, we argue that the standard image editing dimensions have potential to be mapped to the T2I reward space: image quality can transfer directly, prompt following can be aligned through a description of the desired visual state, and reference consistency admits a coarse semantic conversion by encoding the source content to preserve. However, editing instructions specify relative changes, whereas T2I rewards require self-contained target descriptions; moreover, semantically valid captions from generic vision-language models may be incompatible with the frozen reward. Hence, we further introduce Lever-Edit, a two-stage framework that learns a reward-aligned captioner for counterfactual target descriptions, freezes it, and optimizes the editing policy solely with the transferred T2I reward. Experiments show competitive editing alignment and source preservation against editing-reward-based fine-tuning, while outperforming intuitive transfer baselines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑