arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

评估评估者:诊断大型多模态模型的AI生成图像评估能力

Evaluating the Evaluators: Diagnosing Large Multimodal Models for AI-Generated Image Assessment

Yu Zhao, Jiarui Wang, Huiyu Duan, Ye Zhao, Jutao Tang, Juntong Wang, Guangtao Zhai, Xiongkuo Min

arXiv 2609.37576首次发表:更新:

发表机构

Shanghai Jiao Tong University; Dalian University of Technology(上海交通大学; 大连理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对AI生成图像评估缺乏统一基准的问题,提出SQUARE-Bench,从语义、质量、真实性、责任四方面系统评估23个大型多模态模型,并验证其可用于引导迭代编辑。

AI 中文摘要

随着文本到图像(T2I)生成技术的快速发展,稳健的评估变得至关重要且充满挑战,因为传统指标无法捕捉细粒度的对齐性和生成伪影。尽管大型多模态模型(LMMs)越来越多地被用作评估者,现有基准通常孤立地研究语义理解、质量感知和真实性识别,而很大程度上忽视了责任检测。这留下了统一和全面验证的空白。为填补这一空白,我们引入了SQUARE-Bench,一个全面的基准,系统地评估LMM作为AI生成图像评估者在四个方面的能力:语义、质量、真实性和责任。SQUARE-Bench引入了一个包含38个子维度的细粒度分类体系,用于评估从22个不同模型(从传统到最先进的生成器)中采样的近10K张AI生成图像,并辅以超过3K张真实世界图像。这些图像配有精心设计的问题-答案对。对23个LMM的大量实验表明,顶级专有模型,如Gemini-3-Pro,已经超过了单个人类专家基线。然而,模型之间的性能差距仍然显著,在细粒度推理和特定领域鲁棒性方面表现出显著差异。除基准测试外,我们进行了LMM引导的迭代编辑的概念验证研究,其中特定维度的LMM向固定图像编辑器提供诊断反馈。由此产生的引导系统在语义、真实性和责任方面产生了选择性改进,同时表现出一致的视觉质量权衡。SQUARE-Bench既可以作为表征LMM评估者能力的诊断工具,也可以用于研究它们在T2I生成细化中的应用。该基准和数据集将在发表后发布。

英文摘要

With the rapid advancement of text-to-image (T2I) generation, robust evaluation becomes critical yet challenging, as traditional metrics fail to capture fine-grained alignment and generative artifacts. While large multimodal models (LMMs) are increasingly adopted as evaluators, existing benchmarks typically study semantic understanding, quality perception, and authenticity identification in isolation, while largely neglecting responsibility detection. This leaves a gap in unified and comprehensive validation. To bridge this gap, we introduce SQUARE-Bench, a comprehensive benchmark that systematically evaluates LMM capabilities as evaluators of AI-generated images across four aspects: Semantics, Quality, Authenticity, and Responsibility. SQUARE-Bench introduces a granular taxonomy of 38 sub-dimensions to evaluate nearly 10K AI-generated images sampled from 22 diverse models, ranging from legacy to state-of-the-art generators, complemented by over 3K real-world images. The images are annotated with curated question-answering pairs. Extensive experiments on 23 LMMs reveal that top proprietary models, such as Gemini-3-Pro, already outperform the individual human expert baseline. However, the performance gap between models remains significant, exhibiting notable disparities in fine-grained inference and domain-specific robustness. Beyond benchmarking, we conduct a proof-of-concept study of LMM-guided iterative editing, in which dimension-specific LMMs provide diagnostic feedback to fixed image editors. The resulting guided system yields selective improvements in semantics, authenticity, and responsibility, while exhibiting a consistent visual-quality trade-off. SQUARE-Bench can serve as both a diagnostic tool for characterizing LMM evaluator capabilities and studying their use in T2I generation refinement. The benchmark and dataset will be released upon publication.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑