发表机构
SAP; College of Computing and Data Science, Nanyang Technological University; Microsoft(思爱普; 南洋理工大学计算与数据科学学院; 微软公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究发现VLMs的视觉注意力忠实性具有异质性,分三种模式,人工标注区域与模型注意力存在约60%的全面性契合度,且该特性随任务和模型架构变化。
AI 中文摘要
在自然语言处理(NLP)领域,注意力权重是否忠实地反映模型推理过程已引发广泛讨论,但对于视觉-语言模型(Vision-Language Models, VLMs)的视觉模态,该问题仍未得到充分探索。我们通过对当前VLMs进行因果扰动分析来填补这一空白,评估按注意力排序的视觉标记的全面性与充分性差距。分析表明,视觉注意力的忠实性具有异质性,体现为三种不同的处理模式:忠实-充分模式,即前k个注意力标记对预测而言既是必要的也是充分的;忠实-分布式模式,即前k个标记是必要的,但仍需要更广泛的视觉上下文;非聚焦模式,即不存在任何局部注意力区域是单独必要的,而视觉信息仍是预测的必要触发因素。此外,与模型注意力排序相比,人工标注的真实区域仅在约60%的案例中满足全面性,揭示了模型视觉依赖与人类直觉之间存在系统性差异。我们在VQAv2上的通用视觉问答(VQA)任务以及VRDU和ChartQA上的文档任务中验证了这些模式,表明视觉注意力的忠实性会随处理需求和模型架构发生系统性变化,而非呈现统一的忠实或不忠实状态。
英文摘要
Whether attention weights faithfully reflect model reasoning has been actively debated in NLP, yet this question remains largely unexplored for the visual modality in Vision-Language Models (VLMs). We address this gap through causal perturbation analysis on current VLMs, evaluating both the comprehensiveness and sufficiency gap of attention-ranked visual tokens. Our analysis reveals that visual attention faithfulness is heterogeneous, manifesting in three distinct processing modes: Faithful-Sufficient, where top-$k$ attention tokens are both necessary and sufficient for prediction; Faithful-Distributed, where they are necessary but broader visual context remains required; and Non-Focal, where no localized attention region is individually necessary while visual information remains an essential trigger for prediction. Furthermore, human-annotated ground-truth regions satisfy comprehensiveness in only $\sim 60$% of cases compared with model attention rankings, revealing systematic divergence between model visual reliance and human intuition. We demonstrate these patterns across both general VQA on VQAv2 and document tasks on VRDU and ChartQA, showing that visual attention faithfulness varies systematically with processing demands and model architectures rather than being uniformly faithful or unfaithful.
CommentsEMNLP 2026