发表机构
University of Sydney; Microsoft Research; Fudan University; Nankai University; Shanghai Jiao Tong University; Tsinghua University; University of Science and Technology of China; University of Waterloo(悉尼大学; 微软研究院; 复旦大学; 南开大学; 上海交通大学; 清华大学; 中国科学技术大学; 滑铁卢大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
VibeEdit是一种基于画布指令的图像编辑方法,通过空间标记和注释指定编辑内容,经Qwen-Image-Edit模型训练后,在含419个案例的基准上优于文本引导基线FireRed。
AI 中文摘要
在文本引导的图像编辑中,描述期望的修改通常很直接,但识别目标对象或区域却很麻烦,尤其是当多个对象外观相似时。我们引入一种新的图像编辑界面,允许用户直接在图像上放置空间标记和可选的简短注释,这些注释共同构成画布指令,用于指定编辑位置和修改内容。我们的编辑器VibeEdit遵循这些指令执行对象添加、移除、替换、属性修改和移动操作,无需单独的文本提示。我们构建了155万对带有对象掩码和结构化编辑描述的源-目标编辑对,训练期间从中渲染画布指令。我们通过层分离条件调整Qwen-Image-Edit,该条件分别对源图像和画布指令进行编码以用于图像编辑。我们使用区域加权监督微调训练模型,随后进行 rubric 引导的强化学习,以提高编辑完成度、局部编辑质量和未编辑区域的保留效果。我们在独立构建的、人工筛选的包含419个案例的基准上评估VibeEdit,该基准强调相似对象间的目标选择。VibeEdit的VLM rubric得分为79.9,外部区域PSNR为32.8 dB,而评估中得分最高的文本引导基线FireRed的对应值分别为67.4和24.0 dB。
英文摘要
In text-guided image editing, describing the desired change is often straightforward, but identifying the intended object or region can be cumbersome, especially when several objects look alike. We introduce a new image editing interface that lets users place spatial marks and optional short notes directly on the image. Together, these annotations form a canvas instruction that specifies where to edit and what to change. Our editor, VibeEdit, follows these instructions to perform object addition, removal, replacement, attribute modification, and movement without a separate text prompt. We construct 1.55 million source-target edit pairs with object masks and structured edit descriptions, from which we render canvas instructions during training. We adapt Qwen-Image-Edit with layer-decoupled conditioning that separately encodes source images and canvas instructions for image editing. We train the model with region-weighted supervised fine-tuning, followed by rubric-guided reinforcement learning to improve edit completion, local edit quality, and preservation of unedited regions. We evaluate VibeEdit on an independently constructed, human-curated benchmark of 419 cases emphasizing target selection among similar objects. VibeEdit achieves a VLM rubric score of 79.9 and an outside-region PSNR of 32.8 dB, compared with 67.4 and 24.0 dB for FireRed, the highest-scoring text-instructed baseline in our evaluation.
CommentsProject page: https://zhaojingjing713.github.io/VibeEdit/