arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2503.09837cs.CVcs.AIcs.CL

视觉-语言模型在理解图像变换方面的局限性

On the Limitations of Vision-Language Models in Understanding Image Transforms

  • Cohere for AI Community(Cohere AI社区)
  • Arbisoft(Arbisoft公司)
  • Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Ahmad Mustafa Anis, Hasnain Ali, Saquib Sarfraz

更新

AI总结:

本文通过构建增强版Flickr8k数据集评估CLIP与SigLIP,揭示VLM对图像变换理解不足,并分析其对图像编辑等下游任务及Image2Image模型的影响。

AI中文摘要:

视觉-语言模型(VLM)已在图像/视频生成、视觉问答、多模态聊天机器人和视频理解等各类下游任务中展现出巨大潜力。然而,这些模型往往难以处理基本的图像变换。本文研究了VLM在图像层面的理解能力,具体考察OpenAI的CLIP和Google的SigLIP。研究结果表明,这些模型缺乏对多种图像级增强的理解。为推动该研究,我们创建了Flickr8k数据集的增强版本,将每张图像与所应用变换的详细描述配对。我们进一步探讨了这一缺陷如何影响下游任务,尤其是图像编辑,并评估了最先进的Image2Image模型在简单变换上的性能。

英文摘要:

Vision Language Models (VLMs) have demonstrated significant potential in various downstream tasks, including Image/Video Generation, Visual Question Answering, Multimodal Chatbots, and Video Understanding. However, these models often struggle with basic image transformations. This paper investigates the image-level understanding of VLMs, specifically CLIP by OpenAI and SigLIP by Google. Our findings reveal that these models lack comprehension of multiple image-level augmentations. To facilitate this study, we created an augmented version of the Flickr8k dataset, pairing each image with a detailed description of the applied transformation. We further explore how this deficiency impacts downstream tasks, particularly in image editing, and evaluate the performance of state-of-the-art Image2Image models on simple transformations.

补充信息

↑