arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05149cs.CLcs.CV

从视觉到语言:探究多模态决策中的因果信息流

From Vision to Language: Investigating Causal Information Flow in Multimodal Decision-Making

  • Fondazione Bruno Kessler (FBK)(布鲁诺·凯塞勒基金会)
  • Università di Roma La Sapienza(罗马第一大学)
  • Utrecht University(乌得勒支大学)
  • University of Pisa(比萨大学)

机构由 AI 辅助整理,请以论文原文为准。

Davide Testa, Hugh Mee Wong, Alessandro Lenci, Bernardo Magnini, Albert Gatt

AI总结:

本研究在类视频生成式多选任务中,通过分层因果干预探究VLMs的跨模态信息流,明确了视觉信息整合时机、名词与动词的作用及时间推理的特性。

AI中文摘要:

视觉-语言模型(Vision-Language Models, VLMs)通常通过最终预测进行评估,但要理解这些决策是否基于视觉证据,需追踪视觉信息如何作用于基于语言的决策。为此,我们在类视频生成式多选任务中,通过对视频-文本注意力路径施加分层因果干预,探究跨模态信息流,重点针对空间、因果和时间视觉推理。结果表明,视觉信息主要在模型处理候选答案选项时被整合,这些选项是最终决策的主要文本依据;名词在多模态增强中作为语义锚点发挥重要作用,动词则与时间关系处理更相关。最后,我们发现时间推理存在独特模式:VLMs难以重构视频帧间的序列信息,但这种脆弱性也可能反映了定义场景内事件关系的特定时间表达所关联的语言偏差。

英文摘要:

Vision-Language Models are commonly evaluated through their final predictions, but understanding whether these decisions are grounded in visual evidence requires tracing how visual information contributes to language-based decisions. With this purpose in mind, we investigate cross-modal information flow in a video-based generative multiple-choice-like setting by applying a layer-wise causal intervention on video-text attention pathways. We target spatial, causal, and temporal visual reasoning. Our results show that visual information is mainly integrated while the model processes the candidate answer options, which serve as the primary textual grounding sites for the final decision. We further show that nouns play an important role as semantic anchors during multimodal enrichment, while verbs are more relevant when temporal relations are processed. Finally, we identify a distinct pattern in temporal reasoning, suggesting that VLMs struggle to reconstruct sequential information across video frames, but we remark that such fragility may also reflect linguistic biases associated with specific temporal expressions used for defining the relation between events within a scene.

补充信息

↑