arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于一致多参考图像编辑的评估-验证奖励

Evaluation-Verification Reward for Consistent Multi-Reference Image Editing

Yingmao Miao, Pengfei Zhang, Xiaochen Lv, Meng Yu, Lei Sun, Xiangxiang Chu, Chao Shen, Chenhao Lin

arXiv 2607.29025首次发表:更新:

发表机构

Xi’an Jiaotong University; Amap, Alibaba; Shanghai Jiao Tong University(西安交通大学; 高德,阿里巴巴; 上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多参考图像编辑的视觉一致性与奖励模型缺失问题,提出多维度评估-验证奖励EVR,实现无需架构改动的现成编辑器RL微调,性能优于Qwen-Image-Edit并达到或超NanoBanana。

AI 中文摘要

尽管近期图像编辑模型已取得快速进展,但多参考编辑仍具挑战性,尤其难以保持跨参考的视觉一致性和整体视觉协调性。强化学习在文本到图像生成和单图像编辑中已被证明非常有效,但其向多参考编辑的扩展因缺乏能捕捉多图像关系约束的合适奖励模型而受阻。此外,直接使用多模态大语言模型(MLLM)作为零样本评估器面临一个关键矛盾:易产生幻觉的长形式推理与短形式判断有限的演绎能力之间的矛盾。我们提出一种多维度评估-验证奖励(EVR)来解决这些问题。EVR将评估分解为不同的视觉标准;对于每个标准,MLLM评估器生成多个候选假设,验证器将每个主张锚定到具体视觉证据以接受或拒绝,从而产生可靠且细粒度的奖励信号。结合可扩展的数据管道,我们的方法无需架构改动即可实现现成编辑器的强化学习微调。大量实验表明,该方法在基础Qwen-Image-Edit上取得了显著提升,使一致性和协调性达到或超过NanoBanana。

英文摘要

While recent image editing models have made rapid progress, multi-reference editing remains challenging, particularly in maintaining visual consistency across references and ensuring overall visual harmony. Reinforcement learning has proven highly effective for text-to-image generation and single-image editing, but its extension to multi-reference editing is hindered by the absence of suitable reward models that capture multi-image relational constraints. Moreover, naively using multimodal large language models(MLLMs) as zero-shot evaluators faces a key tension between hallucination-prone long-form reasoning and the limited deductive power of short-form judgments. We address these issues with a Multi-dimensional Evaluation-Verification Reward(EVR). EVR decomposes evaluation into distinct visual criteria; for each criterion, an MLLM Evaluator generates multiple candidate hypotheses, and a Verifier grounds each claim in concrete visual evidence to accept or reject it, producing reliable and fine-grained reward signals. Together with a scalable data pipeline, our method enables RL fine-tuning of off-the-shelf editors without architectural changes. Extensive experiments show substantial gains over the base Qwen-Image-Edit, improving consistency and harmony to match or surpass NanoBanana.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑