arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.06021cs.CVcs.AI

可扩展的最小变更学习用于可控图像编辑

Scalable Minimal-Change Learning for Controllable Image Editing

  • Sydney AI Centre, The University of Sydney(悉尼大学悉尼人工智能中心)
  • University of Melbourne(墨尔本大学)

机构由 AI 辅助整理,请以论文原文为准。

Shuo Chen, Fengming Huang, Yu Yao, Mingming Gong, Tongliang Liu

AI总结:

针对图像编辑中非预期更改问题,提出基于强化学习和智能体视觉语言奖励模型的最小变更优化方法ARRO,显著提升编辑精度并减少偏离目标变化。

AI中文摘要:

图像编辑应仅改变指令指定的属性,同时保留其他所有内容,然而当前方法常常产生非预期的更改。我们将这一最小变更原则视为基于指令编辑的优化目标。潜在L1正则化在现代非线性生成器中是输出局部性的较差代理,且通常需要无法大规模获得的监督。我们转而使用强化学习优化编辑结果。一个智能体视觉语言奖励模型审计每个源图像、指令和编辑图像,针对两种失败类型:未执行的请求变更和意外变更。一个分组级评分标准合并并验证这些问题,以在候选编辑之间提供一致的奖励,无需逐指令的人工标注。在FLUX.1 Kontext-dev上,ARRO将MinEval、MagicBrush、AnyBench和Emu-Edit的平均EditScore从5.21提升至5.88。在600个评估示例中,相对于基础编辑器,它减少了8.4%的偏离目标像素变化。奖励和SFT对照、盲人评估以及向OmniGen2的迁移提供了补充证据。代码:此https URL

英文摘要:

Image editing should change only the attributes specified by an instruction while preserving everything else, yet current methods often make unintended changes. We treat this minimal-change principle as an optimization objective for instruction-based editing. Latent L1 regularization is a poor proxy for output locality in modern nonlinear generators and often requires supervision unavailable at scale. We instead optimize edit outcomes with reinforcement learning. An agentic vision-language reward model audits each source image, instruction, and edited image for two failure types: unimplemented requested changes and unintended changes. A group-level rubric merges and verifies these issues to provide consistent rewards across candidate edits without per-instruction human annotations. On FLUX.1 Kontext-dev, ARRO raises average EditScore from 5.21 to 5.88 across MinEval, MagicBrush, AnyBench, and Emu-Edit. On 600 evaluation examples, it reduces off-target pixel change by 8.4% relative to the base editor. Reward and SFT controls, blinded human evaluations, and transfer to OmniGen2 provide complementary evidence. Code: https://github.com/Showwwwwwwww/ARRO

补充信息

↑