arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MIEScore:面向多源图像编辑的人类对齐评估方法

MIEScore: Human-Aligned Evaluation for Multi-Source Image Editing

Zitong Xu, Huiyu Duan, Xinyun Zhang, Weifei Xiong, Tianyi Zheng, Xiongkuo Min, Qiang Hu, Zhengxue Cheng, Bo Li, Guangtao Zhai

arXiv 2608.02059首次发表:更新:

AI 中文总结

针对多源图像编辑(MIE)缺乏人类对齐评估基准的问题,研究构建了首个MIE基准MIE-Bench,并提出基于MLLM的评估模型MIEScore,其对齐人类偏好性能最优且泛化性良好。

AI 中文摘要

近期统一多模态模型的进展大幅提升了文本引导的图像编辑能力,尤其是Nano-Banana-Pro、GPT-Image-2等模型展现出多源图像编辑(MIE)的新兴能力,涵盖对象合成、人物-背景组合、跨图像风格融合等任务。但现有基准和图像编辑评估(IEQA)方法主要聚焦单图像编辑任务,严重忽视了更具挑战性的MIE场景,因此亟需一套全面的、人类对齐的MIE基准。为此,我们推出首个具备细粒度人类偏好标注的大规模多图像编辑基准MIE-Bench,包含16个任务的3000个编辑实例,每个实例涉及两张以上源图像及编辑提示,还有12个前沿编辑模型生成的36000张编辑图像,以及覆盖视觉质量、指令遵循、属性保留维度的超108000个平均意见得分(MOS)。基于MIE-Bench,我们提出MIEScore,这是一种经技能优化和多维度监督微调增强的、基于多模态大语言模型(MLLM)的评估模型,旨在为MIE提供人类对齐的反馈。大量实验表明,MIEScore在与人类偏好对齐方面达到了当前最优性能,且在其他IEQA数据集上具备良好泛化能力,数据集和模型可在指定网址获取。

英文摘要

Recent advances in unified multimodal models have significantly improved text-guided image editing abilities. In particular, models such as Nano-Banana-Pro and GPT-Image-2 demonstrate emerging capabilities in multi-source image editing (MIE), including tasks such as object synthesis, person-background composition, and cross-image style fusion. However, existing benchmarks and image editing assessment (IEQA) methods remain primarily focused on single-image editing tasks and largely overlook the more challenging setting of MIE. This highlights the urgent need for a comprehensive and human-aligned benchmark for MIE. To this end, we introduce MIE-Bench, the first large-scale multiple image editing benchmark with fine-grained human preference annotations. Specifically, MIE-Bench includes 3,000 editing instances across 16 tasks, each involving more than two source images and an editing prompt, together with 36K edited images produced by 12 state-of-the-art editing models and over 108K mean opinion scores (MOSs) covering visual quality, instruction following, and attribute preservation. Based on MIE-Bench, we propose MIEScore, a multimodal large language model (MLLM)-based evaluation model enhanced with skill optimization and multi-dimensional supervised fine-tuning, to provide human-aligned feedback for MIE. Extensive experiments show that MIEScore achieves state-of-the-art performance in aligning with human preferences and generalizes well across other IEQA datasets. Both the dataset and the model are available at https://github.com/IntMeGroup/MIEScore.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑