arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CPI-Bench:面向真实世界图像编辑的综合、实用且智能的基准测试

CPI-Bench: A Comprehensive, Practical and Intelligent Benchmark for Real-World Image Editing

Qinye Zhou, Jun Zheng, Yongchao Du, Yuan Wang, Zhengrui Chen, Zuan Gao, Taihang Hu, Chao Lin, Yefeng Shen, Xingjian Wang, Zhao Wang, Zhengtao Wu, Xiaoli Xu, Zhengze Xu, Hao Yan, Denghui Yang, Yuhang Yu, Huayu Zhang, Mingzhou Zhang, Mengting Chen

arXiv 2608.14546首次发表:更新:

发表机构

Alibaba Group(阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对现有图像编辑基准测试的局限,提出CPI-Bench这一综合实用的智能基准,其含三个核心子集,可有效区分模型性能,与人类评估偏好高度对齐,为模型优化提供指导。

AI 中文摘要

随着图像编辑模型的快速发展及其在各领域的广泛应用,将这些模型能力直接部署到真实场景中的需求日益迫切。然而,现有基准测试仍局限于简单的单图像任务,覆盖维度有限,无法有效区分不同模型的性能,因此无法可靠评估模型在复杂多图像编辑、高要求推理指令及实际部署场景中的表现。为解决这些局限,我们提出CPI-Bench,这是一个面向真实世界图像编辑的综合、实用且智能的基准测试。CPI-Bench包含三个核心子集:CPI-General-Bench,它全面覆盖各类编辑任务,首次纳入多图像编辑评估;CPI-Practical-Bench,聚焦高频真实用户应用场景;CPI-Intelligent-Bench,专注评估高要求的基于推理的编辑能力。基于CPI-Bench对主流图像编辑模型的评估结果表明,该基准测试增强了模型间的性能区分度,全面且可靠地量化了模型在通用编辑能力、实际部署效能及高级推理编辑方面的差距,为未来图像编辑模型的优化提供了宝贵指导。关键的是,我们的排名分析显示,CPI-Bench与Arena Image Edit Leaderboard的对齐度最高,表明它忠实地捕捉了人类评估者的偏好和感知判断,可作为真实用户体验的可靠替代指标。

英文摘要

With the rapid advancement of image editing models and their widespread application across various domains, there is an increasingly urgent need to deploy these model capabilities directly into real-world scenarios. However, existing benchmarks remain confined to simple single-image tasks, suffering from limited coverage dimensions and an inability to effectively differentiate performance among diverse models. Consequently, they fail to reliably evaluate model performance in complex multi-image editing, highly demanding reasoning instructions, and practical deployment settings. To address these limitations, we propose CPI-Bench, a Comprehensive, Practical and Intelligent benchmark for real-world image editing. CPI-Bench comprises three core subsets: CPI-General-Bench, which comprehensively covers diverse editing tasks and introduces multi-image editing evaluation; CPI-Practical-Bench, which focuses on high-frequency real-user application scenarios; and CPI-Intelligent-Bench, which is dedicated to evaluating capabilities in highly demanding reasoning-based editing. Evaluation results of mainstream image editing models based on CPI-Bench demonstrate that CPI-Bench enhances performance differentiation among models. It provides a comprehensive and reliable quantification of gaps in general editing capabilities, practical deployment efficacy, and advanced reasoning-based editing, offering invaluable guidance for the future optimization of image editing models. Crucially, our ranking analysis reveals that CPI-Bench achieves the highest alignment with the Arena Image Edit Leaderboard, indicating stronger consistency with public human preference rankings, serving as an effective proxy for public human evaluations.

Comments13 pages, benchmark report

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑