arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉输入及其构图方式影响大型视觉-语言模型生成的基于属性的描述

Visual Input and Its Framing Affect Attribute-based Descriptions Produced by Large Vision-Language Models

Xiaomeng Wang, Martha Larson, Zhengyu Zhao

arXiv 2609.18345首次发表:更新:

AI 中文总结

本文发现大型视觉-语言模型在存在图像时,即使文本提示不针对具体实例,响应也会受图像及其构图影响,并量化了物理术语占比的变化,强调评估鲁棒性需考虑视觉输入。

AI 中文摘要

大型视觉-语言模型(LVLMs)通常仅使用单个文本提示作为输入,或再加上一张图像。在本文中,我们证明当图像存在时,即使文本提示并非针对该图像中的特定实例(而仅针对其所属的概念),响应仍会受到影响。例如,当文本提示仅要求描述某个犬种的属性时,一张描绘该犬种中特定犬只的图像会使响应发生偏移。此外,该特定实例在图像中的构图方式将决定响应向哪个方向偏移。详细分析还表明,在响应中,物理术语的占比从仅文本时的18%增加到以主体为中心(主体在情境中)构图时的45%(40%)。总体而言,视觉线索对LVLMs的意外影响凸显了在评估LVLMs鲁棒性时,理解图像的存在及其构图方式的必要性。

英文摘要

Large vision-language models (LVLMs) are commonly used with only a single text prompt as the input, or plus an image. In this paper, we demonstrate that when the image exists, even if the text prompt is not about the specific instance (but only the concept it belongs to) in that image, the response would still be affected. For example, when the text prompt only asks for the attribute descriptions of a dog breed, an image depicting a specific dog from that breed would shift the response. Further, how the specific instance is framed in that image would determine towards which the response shifts. Detailed analyses also reveal that in the response, physical terms increase from 18% for text-only to 45% (40%) for subject-focused (subject-in-situation) framings. Overall, the unexpected effects of visual cues on LVLMs highlight the need to understand the presence of an image and its framing when evaluating the robustness of LVLMs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑