发表机构
ETH Zurich(苏黎世联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出干预性协议评估视觉-语言模型对模态缺失的自解释,发现模型系统性高估可用证据充分性并低估恢复缺失模态的影响,建议以可执行干预作为行为真值。
AI 中文摘要
视觉-语言模型越来越多地被用于某些输入模态可能不可用的场景,然而我们对其能否忠实解释这种缺失信息如何影响自身预测知之甚少。我们提出了一种用于评估模态动态自解释的干预性协议:模型需陈述每个模态单独能支持什么、恢复缺失模态是否会改变其答案,以及现有证据是否充分;随后我们执行相应的模态干预,并将这些声明与模型的实际行为进行比较。我们评估了来自两个模型家族的八个开放权重视觉-语言模型,涵盖四个任务,包括互补性和同构性文本-图像设置以及一个多视角驾驶设置。我们发现模型存在系统性高估现有模态证据充分性的倾向。模型显著低估了恢复缺失模态的影响:任务级中位预测变化率最高为8.8%,而相应的实际执行变化率高达72.1%,在64个模型-任务-条件设置中有62个出现低估。不充分性声明很少见,但一旦产生则很精确:在标记案例中,恢复模态改变答案的中位比例为78-100%。回顾性自解释也表现出相同倾向:在互补性数据上,模型过度归因于单模态充分性;在同构性数据上,模型相对于其实际行为过度归因于单一表示充分性。综合来看,这些结果表明视觉-语言模型系统性地错误描述了其预测对可用和缺失模态证据的依赖,这促使将可执行干预作为评估多模态自解释的行为真值。
英文摘要
Vision-language models (VLMs) are increasingly used in settings where some input modalities may be unavailable, yet we know little about whether they can faithfully explain how such missing information affects their own predictions. We introduce an interventional protocol for evaluating self-explanations of modality dynamics: models state what each modality alone would support, whether restoring a missing modality would change their answer, and whether the available evidence is sufficient; we then execute the corresponding intervention and compare these claims with realized behavior. We evaluate ten VLMs spanning open-weight and proprietary models across four tasks covering mixed, redundant, and unique modality regimes. We find a systematic tendency to overstate the sufficiency of available evidence. Models substantially underestimate the effect of restoring missing modalities: executed change exceeds predicted change in 78 of 80 model-task-condition settings, with task-level median executed change rates reaching 70.1\% while median predicted rates remain at most 9.6\%. Insufficiency claims have low recall, leaving many cases in which behavior changes despite a stated claim of sufficiency. Retrospective self-explanations show the same tendency, over-crediting single-input sufficiency in mixed regimes and interchangeability in redundant ones. Together, these results show that VLMs systematically mischaracterize how their predictions depend on available and missing evidence, motivating executable interventions as a behavioral test of multimodal self-explanations.
CommentsAccepted at VLM4RWD at NeurIPS 2026