arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉不敏感差距:诊断视觉-语言模型何时无法使用视觉证据

The Visual Insensitivity Gap: Diagnosing When Vision-Language Models Fail to Use Visual Evidence

Genpei Zhang

arXiv 2609.00868首次发表:更新:

发表机构

University of Wisconsin–Madison(威斯康星大学麦迪逊分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究诊断发现视觉-语言模型存在无法利用视觉证据的视觉不敏感差距,提出VSI量化该差距,明确其适用场景,表明其可作为条件集成组件使用。

AI 中文摘要

视觉-语言模型(VLMs)通过多模态基准上的总体准确率进行评估,这种做法隐含假设模型会使用其视觉输入。我们表明,在六个VLMs和三个感知基准中,40%至97%的样本上这一假设不成立:模糊与问题相关的视觉区域几乎不会改变下一个词的分布。我们将这一现象命名为视觉不敏感差距,并通过逐样本视觉敏感性指数(VSI)对其进行量化。该差距是样本的属性,而非模型的属性:VSI在不同模型间的排名具有相关性(总体平均斯皮尔曼相关系数ρ=+0.40,置换检验p<10^-3),因此即使是仅共享对比预训练视觉编码器架构的VLMs,也会对相同样本标记为不敏感。其机制具体:在不敏感样本上,对每个模型自身视觉编码器的线性探针区分扰动图像与干净图像的准确率为0.72至0.79,但模型的argmax词仅在2%至11%的相同样本上发生变化,每个模型的编码器-大语言模型(LLM)差距均超过0.65。逐单元映射VSI的诊断效用,可得到强适用场景(能力较强的VLMs上的多选推理:AUROC=0.85至0.87)与弱适用场景(校准良好的事实性任务,其中softmax置信度已足够)。VSI并非通用的最佳弃权(不执行)信号,它是样本固有视觉忽略失败的指标,最适合作为条件集成组件使用。

英文摘要

Vision-language models are evaluated by aggregate accuracy on multimodal benchmarks, a practice that implicitly assumes the model uses its visual input. We show this assumption fails on 40%--97% of samples across six VLMs and three perceptual benchmarks: blurring the question-relevant visual region leaves the next-token distribution nearly unchanged. We name this phenomenon the Visual Insensitivity Gap and quantify it with a per-sample Visual Sensitivity Index (VSI). The gap is a property of samples, not of models: VSI ranks correlate across models (grand-mean Spearman rho=+0.40, permutation p<10^-3), so the same samples are flagged insensitive by VLMs sharing no architectural detail beyond a contrastively pretrained vision tower. The mechanism is concrete: on the insensitive samples, a linear probe on each model's own vision tower distinguishes perturbed from clean images at 0.72--0.79 accuracy, yet the model's argmax token changes on only 2%--11% of the same samples, an encoder--LLM gap above 0.65 on every model. Mapping VSI's diagnostic utility cell by cell surfaces a strong regime (multi-choice reasoning on capable VLMs: AUROC=0.85--0.87) and a weak regime (well-calibrated factuality, where softmax confidence already leads). VSI is not a universal best abstention signal; it is a sample-intrinsic indicator of vision-ignoring failure, best used as a conditional ensemble component.

Comments17 pages (7-page main text plus technical appendix), 10 figures, 6 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑