语言上下文对视觉语言模型中的视觉表示进行重编码
Linguistic Context Recodes Visual Representations in Vision-Language Models
浏览论文内容
中文总结 AI 辅助
本研究揭示视觉语言模型中语言上下文会重编码视觉表示,通过对比引导向量和属性调制两种机制,为跨模态动态处理提供了因果证据。
中文摘要 AI 辅助
目标导向的视觉处理是人类视觉智能的标志,其产生的表示可支持分类或搜索等下游任务。尽管视觉语言模型(VLMs)常面临相同任务,但其在呈现目标导向语言时重编码视觉表示的能力却鲜有研究。过往研究多将VLMs中的视觉表示视为由语言表示操控的静态视觉信息存储库。本研究提供了语言诱导视觉表示重编码的两个具体实例证据:其一,我们识别出一种抽象参考表示,它表示自然语言提示下哪些对象与目标相关。我们提取对应该参考表示的对比引导向量,并证明它们对模型预测具有因果影响。这些参考表示具有抽象性,可泛化到不同对象、不同任务情境,甚至从合成图像到自然图像。其二,我们证明语言诱导的属性调制:后期层会选择性放大对象视觉表示中与目标相关的属性,且在一系列不同提示中都能观察到该现象。最后,我们通过因果干预证明,属性调制介导了VLM的响应分布。综上,我们的结果支持VLMs中跨模态处理的更动态解释——视觉令牌并非静态信息存储库,而是被调制以支持语言表述的查询。
英文摘要
Goal-directed visual processing is a hallmark of human visual intelligence, resulting in representations that support downstream tasks such as categorization or search. Though vision-language models (VLMs) are often faced with these same tasks, their ability to recode visual representations when presented with goal-directed language remains poorly characterized. Indeed, prior work largely treats visual representations in VLMs as static repositories of visual information that are manipulated by language representations. In the present work, we provide evidence for two concrete instances of language-induced recoding of visual representations. First, we identify an abstract reference representation that denotes which objects are goal-relevant under a natural language prompt. We extract contrastive steering vectors corresponding to this reference representation and demonstrate that they are causally implicated in model predictions. These reference representations are abstract in that they generalize to different objects, different task contexts, and even from synthetic to naturalistic images. Second, we demonstrate language-induced attribute modulation: later layers selectively amplify goal-relevant attributes in visual representations of objects. We demonstrate this phenomenon across a range of different prompts. Finally, we provide a causal intervention that demonstrates that attribute modulation mediates a VLM's response distribution. Together, our results support a more dynamic account of cross-modality processing in VLMs -- rather than vision tokens serving as static repositories of information, they are modulated to support queries articulated in language.
发表机构
- Brown University(布朗大学)
机构由 AI 辅助整理,请以论文原文为准。