用于解释视觉语言模型中视觉素养的注意力引导显著图
Attention-Guided Saliency Maps for Interpreting Visualization Literacy in VLMs
浏览论文内容
中文总结 AI 辅助
研究视觉语言模型如何解释数据可视化这一问题,提出针对图像文本生成的轻量级诊断显著图方法,通过聚合注意力并映射到图像,生成快速无梯度显著图,用删除度量评估,揭示模型注意力分配情况。
中文摘要 AI 辅助
理解视觉语言模型(VLM)如何解释数据可视化仍是一个未解决的问题,随着这些模型用于需要可靠推理的分析任务,其重要性日益凸显。我们针对使用变压器模型(当前可视化解释中的先进模型)进行图像文本生成引入了一种轻量级诊断显著图方法。我们的方法聚合语言模型在所有头和层上对视觉令牌的注意力,然后将此注意力映射回视觉编码器的补丁网格以在图像上定位,在每个生成的答案令牌与其关注的图像区域之间建立直接对应关系。这产生了快速、无梯度的显著图,揭示了VLM在答案生成过程中如何在视觉元素间分配注意力,可检查模型注意力是否与语义相关组件对齐。我们使用删除度量评估方法,验证显著图对模型行为的因果忠实性。
英文摘要
Understanding how vision-language models (VLMs) interpret data visualizations remains an open problem, and is increasingly important as these models are used for analytical tasks where reliable reasoning is essential. We introduce a lightweight, diagnostic saliency map method tailored for text generation over images using transformer models, the current state-of-the-art models in visualization interpretation. Our approach aggregates the language model's attention over the visual tokens across all heads and layers, then maps this attention back onto the vision encoder's patch grid to localise it over the image, producing a direct correspondence between each generated answer token and the image regions it attended to. This yields fast, gradient-free saliency maps that expose how VLMs allocate focus across visual elements during answer generation, enabling inspection of whether model attention aligns with semantically relevant components. We evaluate our approach using a deletion metric which validates the causal faithfulness of our saliency maps to the model's behavior.
发表机构
- The Alan Turing Institute(艾伦·图灵研究所)
机构由 AI 辅助整理,请以论文原文为准。