arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VibeEdit:基于画布指令的图像编辑

VibeEdit: Image Editing with Canvas Instructions

Jinjing Zhao, Fangyun Wei, Yitong Wang, Xiuyu Wu, Yunuo Chen, Yang Yue, Sirui Zhang, Wenbo Wang, Hongyang Zhang, Dong Chen, Yan Lu, Chang Xu

arXiv 2610.12229首次发表:更新:

发表机构

University of Sydney; Microsoft Research; Fudan University; Nankai University; Shanghai Jiao Tong University; Tsinghua University; University of Science and Technology of China; University of Waterloo(悉尼大学; 微软研究院; 复旦大学; 南开大学; 上海交通大学; 清华大学; 中国科学技术大学; 滑铁卢大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

VibeEdit是一种基于画布指令的图像编辑方法,通过空间标记和注释指定编辑内容,经Qwen-Image-Edit模型训练后,在含419个案例的基准上优于文本引导基线FireRed。

AI 中文摘要

在文本引导的图像编辑中,描述期望的修改通常很直接,但识别目标对象或区域却很麻烦,尤其是当多个对象外观相似时。我们引入一种新的图像编辑界面,允许用户直接在图像上放置空间标记和可选的简短注释,这些注释共同构成画布指令,用于指定编辑位置和修改内容。我们的编辑器VibeEdit遵循这些指令执行对象添加、移除、替换、属性修改和移动操作,无需单独的文本提示。我们构建了155万对带有对象掩码和结构化编辑描述的源-目标编辑对,训练期间从中渲染画布指令。我们通过层分离条件调整Qwen-Image-Edit,该条件分别对源图像和画布指令进行编码以用于图像编辑。我们使用区域加权监督微调训练模型,随后进行 rubric 引导的强化学习,以提高编辑完成度、局部编辑质量和未编辑区域的保留效果。我们在独立构建的、人工筛选的包含419个案例的基准上评估VibeEdit,该基准强调相似对象间的目标选择。VibeEdit的VLM rubric得分为79.9,外部区域PSNR为32.8 dB,而评估中得分最高的文本引导基线FireRed的对应值分别为67.4和24.0 dB。

英文摘要

In text-guided image editing, describing the desired change is often straightforward, but identifying the intended object or region can be cumbersome, especially when several objects look alike. We introduce a new image editing interface that lets users place spatial marks and optional short notes directly on the image. Together, these annotations form a canvas instruction that specifies where to edit and what to change. Our editor, VibeEdit, follows these instructions to perform object addition, removal, replacement, attribute modification, and movement without a separate text prompt. We construct 1.55 million source-target edit pairs with object masks and structured edit descriptions, from which we render canvas instructions during training. We adapt Qwen-Image-Edit with layer-decoupled conditioning that separately encodes source images and canvas instructions for image editing. We train the model with region-weighted supervised fine-tuning, followed by rubric-guided reinforcement learning to improve edit completion, local edit quality, and preservation of unedited regions. We evaluate VibeEdit on an independently constructed, human-curated benchmark of 419 cases emphasizing target selection among similar objects. VibeEdit achieves a VLM rubric score of 79.9 and an outside-region PSNR of 32.8 dB, compared with 67.4 and 24.0 dB for FireRed, the highest-scoring text-instructed baseline in our evaluation.

CommentsProject page: https://zhaojingjing713.github.io/VibeEdit/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑