arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AnswerMap: 从答案后验中获取VLM的忠实空间可解释性

AnswerMap: Faithful Spatial Interpretability of VLMs from Answer Posteriors

Mohamed Eltahir, Fardows Adam, Duaa M. Tahir, Lama Alamoudi, Sana Ammar, Atheer A. Alboloshi, Jory Albluey, Tanveer Hussain, Naeemullah Khan

arXiv 2609.35247首次发表:更新:

发表机构

King Abdullah University of Science and Technology (KAUST); Edge Hill University(阿卜杜拉国王科技大学; 埃奇希尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

AnswerMap从输出头构建无需训练的黑盒空间图,通过行/列后验外积生成查询条件图,在四个模型上验证了忠实性(AUC 0.85),并通过读出支持定位、幻觉检测和答案修正。

AI 中文摘要

当视觉语言模型(VLM)回答视觉查询时,当前的可解释性工具依赖于文本理由(使用不匹配的模态)或内部读出(产生过早,无法反映最终输出,且需要白盒访问模型)。我们引入了AnswerMap,一种无需训练、任务无关、黑盒的视觉理由,由输出头构建。图像被切割成K行和K列条带,每个条带单独与查询以是/否相关性问题的形式呈现给冻结模型。行和列“是”后验的外积给出查询条件下的空间图。关键在于,通过在AnswerMap之上定义固定的读出R(例如,期望、最大值),我们可以原生地推导出连续输出(如位置)。这绕过了连续输出任务对离散文本标记的依赖,并保证了图像依赖的答案。然而,理由可能被虚构,因此我们通过两个测试在四个模型和三个查询分布上验证AnswerMap:(a)与模型自身生成的点的一致性,(b)删除图区域的影响。该图落在模型指向的位置(AUC 0.85,而注意力为0.38),删除其区域使53%的正确回答翻转(而注意力为19%)。除了建立忠实性,我们通过三个不同的读出展示了图的任务无关实用性:其最大值无需生成即可标记幻觉对象,其期望在模型自身指向失败时正确定位,其顶部质量区域作为裁剪反馈,修复了模型一半的错误答案。因此,AnswerMap为VLM可解释性提供了新视角,并通过其读出为视觉任务提供了超越文本标记的新输出接口。

英文摘要

When a VLM answers a visual query, current interpretability tools rely on text rationales, which use a mismatched modality, or on internal read-outs, which originate too early to reflect the final output and require white-box access to the model. We introduce AnswerMap, a training-free, task-agnostic, black-box visual rationale constructed from the output head. The image is cut into K row and K column bands, each shown alone to the frozen model along with the query in the format of a yes/no relevance question. The outer product of the row and column ``yes'' posteriors gives the query-conditioned spatial map. Crucially, by defining a fixed read-out R (e.g., expectation, maximum) on top of AnswerMap, we can derive continuous outputs like location natively. This bypasses the reliance on discrete text tokens for continuous-output tasks and guarantees an image-dependent answer by construction. However, a rationale can be confabulated, so we validate AnswerMap across four models and three query distributions with two tests: (a) agreement with the model's own generated point and (b) deletion of the map's region. The map lands where the model points (AUC 0.85 against 0.38 for attention), and deleting its region flips 53% of correct answers (against 19% for attention's). Beyond establishing faithfulness, we demonstrate the map's task-agnostic utility through three distinct read-outs: its maximum flags hallucinated objects without generation, its expectation localizes correctly when the model's own pointing fails, and its top-mass region, fed back as a crop, fixes half of the model's wrong answers. AnswerMap thus offers a new lens on VLM interpretability and, through its read-outs, a new output interface for visual tasks beyond text tokens.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑