发表机构
Mohamed bin Zayed University of Artificial Intelligence; Aalto University(穆罕默德·本·扎耶德人工智能大学; 阿尔托大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出Rubric-CEPR自进化图像编辑框架,仅用编辑器自身生成结果,通过规则增强的CEPR奖励验证样本并蒸馏优化,在多数据集上提升了图像编辑性能。
AI 中文摘要
指令引导的图像编辑器已具备很强的能力,但进一步提升仍依赖人工编辑的训练对或外部奖励模型。这类监督信号获取成本高,且可能奖励看似合理的失败结果:一个逼真的输出可能未完成请求的修改,或改动了应保留的内容。本研究仅利用图像编辑器自身的生成结果来改进预训练的图像编辑器,无需人工编辑的目标样本或训练时的外部奖励模型。为此,我们提出名为Rubric-CEPR的自进化框架,该框架通过基于规则增强的对比编辑-保留奖励(Contrastive Edit-Preservation Reward, CEPR),利用编辑器的内部表示来验证其自身样本。具体流程为:规划器从未标注图像中提出结构化编辑指令,编辑器采样多个候选编辑结果,冻结的评价器(Critic)利用编辑器已暴露的特征,通过分解的规则检查(涵盖编辑实现、旧状态移除、内容保留三个维度)对每个候选结果打分;非补偿门(Non-compensatory gates)会拒绝不可行的候选结果,经验证的最优候选结果通过轻量级适配器训练被蒸馏到编辑器中。在Qwen-Image-Edit数据集上,Rubric-CEPR将ImgEdit指标从4.36提升至4.60,增幅达5.5%,在对象隔离任务上提升24.9%,且可迁移至GEdit-Bench和Complex-Edit数据集;该方法还使Step1X-Edit在ImgEdit上提升7.8%。我们希望本方法能成为图像编辑器的坚实基线,这类编辑器可基于自身验证样本实现自我改进,代码已公开于this URL。
英文摘要
Instruction-guided image editors have become highly capable, yet improving them further still depends on human-edited training pairs or external reward models. Such supervision is costly to obtain and can reward plausible failures: a realistic output may leave the requested change undone or alter content that should be preserved. In this work, we strive to improve a pretrained image editor using only its own generations, without human-edited targets or an external training-time reward model. To this end, we propose a self-evolving framework, named Rubric-CEPR, that verifies the editor's own samples with its internal representations through a rubric-augmented Contrastive Edit-Preservation Reward (CEPR). A Planner proposes structured edit instructions from unlabeled images, the Editor samples multiple candidate edits, and a frozen Critic scores each candidate with decomposed rubric checks for edit realization, removal of the old state, and content preservation, using features already exposed by the editor. Non-compensatory gates reject infeasible candidates, and the best verified candidate is distilled into the editor through lightweight adapter training. On Qwen-Image-Edit, Rubric-CEPR improves ImgEdit from 4.36 to 4.60 (+5.5%), with a +24.9% gain on object isolation, and transfers to GEdit-Bench and Complex-Edit. The same procedure also improves Step1X-Edit by +7.8% on ImgEdit. We hope our approach will serve as a solid baseline for image editors that improve themselves from their own verified samples. Our code is publicly available at $\href{https://riteshthawkar.github.io/Rubric-CEPR/}{\text{this URL}}$
CommentsProject Page: $\href{https://riteshthawkar.github.io/Rubric-CEPR/}{\text{this URL}}$