arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VDiff-Bench:一个具有挑战性的细粒度图像差异识别基准

VDiff-Bench: A Challenging Benchmark for Fine-Grained Image Difference Identification

Yixin Wan, Tianle Zheng, Kai-Wei Chang

arXiv 2609.06245首次发表:更新:

发表机构

University of California, Los Angeles(加利福尼亚大学洛杉矶分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

VDiff-Bench是一个包含1,756个四选一问题的基准,用于评估多模态大语言模型在细粒度图像差异识别上的能力,涵盖10个变化类别,实验发现模型在低级变化上表现脆弱。

AI 中文摘要

多模态大语言模型(MLLMs)在通用视觉理解任务(如视觉问答)上表现强劲,但它们常常难以掌握一项基本的比较技能:识别两幅相似图像之间发生了什么变化。我们引入了VDiff-Bench,一个用于细粒度图像差异识别的具有挑战性的多项选择基准。VDiff-Bench包含1,756个基于图像对的四选一问题,涵盖10个变化类别:位置、运动、区域图像颜色、整体图像颜色、出现/消失、噪声/分辨率、纹理、替换/大小、OCR/文本和光照。每个问题对应两个图像输入,并提供4个选项:真实差异、两个困难负样本描述和一个“无差异”干扰项。为了使任务具有挑战性,我们特别策划了基于真实条件的负样本,要求模型从相近的语义候选项中区分出实际变化。对11个最先进的开源和闭源MLLMs进行的实验表明,细粒度视觉比较仍然脆弱:模型在不同来源和变化类别上表现不均,在噪声和纹理等细微低级变化上持续失败。例如,三个7-8B规模的开源MLLMs在语义变化上得分52.5-70.6%,但在噪声和纹理等低级变化上仅得分8.7-33.3%,错误地假设两个图像输入之间没有变化。令人惊讶的是,尽管其他闭源商业模型表现强劲,Grok 4.3在识别图像之间的噪声和纹理差异方面表现出显著的性能下降,明显落后于Kimi K2.5和K3等大型开源模型。总体而言,VDiff-Bench为评估MLLMs中的比较视觉理解提供了一个有针对性的诊断工具,揭示了标准单图像视觉语言任务无法捕获的失败。

英文摘要

Multimodal Large Language Models (MLLMs) perform strongly on general visual understanding tasks such as visual question answering, yet they often struggle with a basic comparative skill: identifying what has changed between two similar images. We introduce VDiff-Bench, a challenging multiple-choice benchmark for fine-grained Image Difference Identification. VDiff-Bench contains 1,756 four-way questions over image pairs and covers 10 change categories: position, motion, regional image color, overall image color, appearance/disappearance, noise/resolution, texture, substitution/size, OCR/text, and illumination. Each question corresponds to two image inputs with 4 choices: the true difference, two hard negative descriptions, and a "no difference" distractor. To make the task challenging, we specifically curate ground-truth-conditioned negatives that require models to distinguish the actual change from nearby semantic alternatives. Experiments with 11 state-of-the-art open- and closed-source MLLMs show that fine-grained visual comparison remains brittle: models exhibit uneven performance across sources and change categories, with persistent failures on subtle low-level changes like noises and textures. For instance, three 7-8B-scale open-source MLLMs score 52.5-70.6% on semantic changes but only 8.7-33.3% on low-level changes like noise and texture, falsely assuming no changes between two image inputs. Surprisingly, despite strong performance of other closed-source commercial models, Grok 4.3 demonstrate remarkable performance drop on identifying noise and texture differences between images, falling significantly behind large open-source models like Kimi K2.5 and K3. Overall, VDiff-Bench provides a targeted diagnostic for evaluating comparative visual understanding in MLLMs, exposing failures that are not captured by standard single-image vision-language tasks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑