发表机构
Singapore Institute of Technology; NVIDIA(新加坡科技学院; 英伟达)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究组合式视觉问答中视觉语言模型失败的机制,引入以操作为中心的框架分解失败模式,揭示四种失败模式及传播路径,表明不同失败类型需不同纠正策略,为提升模型可靠性提供基础。
AI 中文摘要
组合式视觉问答要求视觉语言模型执行多种推理操作,如对象选择、空间关系解析和属性验证。尽管总体性能强劲,但视觉语言模型在此任务上失败的机制基础仍未得到充分探索。为填补这一空白,我们通过研究失败如何与特定推理操作以及它们出现和传播的内部计算路径相关,来分析视觉语言模型中的视觉-操作不对齐。我们引入了一个以操作为中心的机制框架,该框架根据失败产生的推理操作和传播的内部计算路径对视觉语言模型的失败进行分解。我们的分析揭示了四种机制上不同的失败模式:基础失败、推理失败、属性提取失败和语言先验主导失败。每种模式都以视觉基础强度和答案正确性之间的独特关系为特征。通过在所有变压器层应用的三种互补因果干预,我们进一步证明了一种路径分离:基础失败仅通过前馈网络传播,推理失败通过后期层注意力传播,属性提取失败定位到答案位置的前馈计算。这种分离表明不同的失败类型需要根本不同的纠正策略,为有针对性地提高视觉语言模型在多媒体推理中的可靠性提供了原则基础。
英文摘要
Compositional visual question answering requires Vision-Language Models (VLMs) to execute multiple reasoning operations like object selection, spatial relation resolution, and attribute verification. Despite strong aggregate performance, the mechanistic basis of VLM failures on this task remains underexplored. To address this gap, we analyze vision-operation misalignment in VLMs by examining how failures relate to specific reasoning operations and the internal computational pathways through which they arise and propagate. We introduce an Operation-centric mechanistic framework that decomposes VLM failures by both the reasoning operation where they originate and the internal computational pathway through which they propagate. Our analysis reveals four dominant failure modes: grounding failure, reasoning failure, attribute extraction failure, and language-prior dominance, each characterized by a distinct relationship between visual grounding strength and answer correctness. Through three complementary causal interventions applied across all transformer layers, we find that object-selection failures are associated primarily with feedforward computation, multi-step relational failures with late-layer direct attention, and attribute-extraction failures with answer-position feedforward computation. Validation on VSR further shows that single-step spatial failures are concentrated at object-position encoding, distinguishing them from multi-step relational composition. These findings reveal distinct computational bottlenecks across operation types and provide a principled basis for targeted diagnosis of VLM failures in multimedia reasoning.
CommentsAccepted at ACM Multimedia 2026