arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉还是认知?多模态大语言模型的视觉上下文敏感性

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models

Jiaang Li, Chengzu Li, Zhaochong An, Yifei Yuan, Xi Liu, Serge Belongie, Vésteinn Snæbjarnarson

arXiv 2607.26326首次发表:更新:

发表机构

University of Copenhagen; University of Cambridge; ETH Zürich; Clemson University(哥本哈根大学; 剑桥大学; 苏黎世联邦理工学院; 克莱姆森大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究探究多模态大语言模型在视觉证据与先验冲突任务上的失效原因,引入WhatIfVis基准,发现其能编码视觉证据但难以控制对其的依赖,监督微调等方法可提升可控性。

AI 中文摘要

多模态大语言模型(Multimodal Large Language Models, MLLMs)通过将视觉输入与预训练语言模型的丰富先验知识相结合,实现了出色的性能。然而,它们在以视觉为核心的任务上常常失效,尤其是当视觉证据与预训练知识发生冲突时。我们采用两种诊断范式分别探究这些失效原因:(1)通过图像重建探测视觉信息是否可用;(2)测量多模态上下文敏感性,即模型遵循视觉上下文而非语言先验的程度。为支持第二项探究,我们引入了WhatIfVis基准,该基准涵盖五个粗粒度维度(时空、颜色、数量、大小和重量),其问题的答案既可以来自图像,也可以来自先验知识。我们的分析得出三项发现:(i)粗粒度视觉证据得以保留,因为这些属性可从冻结MLLMs的最终层图像token中重建;因此,关于这些属性的问题失效指向的是感知后的利用问题,而非感知过程中视觉编码的退化。(ii)即使被明确指示使用或忽略视觉证据,未在WhatIfVis上进行监督微调(Supervised Fine-Tuning, SFT)的普通模型也表现出不稳定的视觉上下文敏感性。监督微调提升了这种可控性,并能跨领域泛化,而激活修补进一步在所有六个模型的特定架构深度上定位了视觉与先验之间的权衡。(iii)视觉与先验之间的权衡可沿学习到的向量进行控制;应用该引导向量,即使没有任何意图指令,也能提升普通模型的可控性。综上,这些结果重新定位了瓶颈,表明对于我们研究的粗粒度属性,MLLMs会编码视觉证据,但无法可靠地控制对其的依赖。

英文摘要

Multimodal Large Language Models (MLLMs) achieve strong performance by integrating visual inputs with the rich priors of pretrained language models. However, they often fail on vision-centric tasks, especially when visual evidence conflicts with pretrained knowledge. We explore these failures separately using two diagnostic paradigms: (1) probing whether visual information is available, via image reconstruction, and (2) measuring multimodal context sensitivity, the extent to which the model follows visual context versus the language prior. To support the second, we introduce the WhatIfVis, a benchmark spanning five coarse-grained dimensions (spatial-temporal, color, count, size, and weight) whose questions admit answers from either the image or the prior. Our analysis yields three findings: (i) Coarse-grained visual evidence is preserved, as these attributes can be reconstructed from the final-layer image tokens of frozen MLLMs. Failures on questions about these attributes therefore point to post-perceptual utilization, rather than to degraded visual encoding during perception. (ii) Even when explicitly instructed to use or ignore visual evidence, vanilla models (without supervised fine-tuning on the WhatIfVis) show unstable visual context sensitivity. Supervised fine-tuning (SFT) improves this controllability and generalizes across domains, and activation patching further localizes the vision-versus-prior trade-off at architecture-specific depths across all six models. (iii) The vision-versus-prior trade-off is controllable along a learned vector. Applying this steering vector, even without any intent instruction, improves controllability over the vanilla model. Together, these results relocate the bottleneck, indicating that for the coarse attributes we study, MLLMs encode the visual evidence but cannot reliably control their reliance on it.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑