动态线索瓶颈:迈向设计即可解释的视觉问答
Dynamic Clue Bottlenecks: Towards Interpretable-by-Design Visual Question Answering
- University of Pennsylvania(宾夕法尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对多模态大模型视觉问答可解释性不足的问题,提出设计即可解释的动态线索瓶颈模型DCLUB,通过先输出视觉线索再生成答案实现决策可追溯,在推理类问题上优于同类黑盒模型,同时保持主流基准性能。
AI中文摘要:
多模态大语言模型(LLM)的最新进展在视觉问答(VQA)任务中展现出极强的有效性。然而,这类端到端模型的设计本质使其无法为人类所解释,削弱了其在关键领域的可信度与适用性。尽管事后归因解释为理解模型行为提供了一定洞见,但这类解释无法保证忠实于模型的真实决策过程。\n本文针对上述缺陷,提出一种设计即具备可解释性的模型,该模型将决策过程拆解为人类可读的中间解释步骤,让人们能轻松理解模型成败的原因。我们提出动态线索瓶颈模型(Dynamic Clue Bottleneck Model,DCLUB),这是一种面向固有可解释VQA系统设计的方法。DCLUB在VQA决策前提供可解释的中间空间,从设计之初就具备忠实性,同时能保持与黑盒系统相当的性能。\n给定问题时,DCLUB首先返回一组视觉线索:即从图像中提取的视觉显著证据的自然语言表述,随后仅基于这些视觉线索生成输出答案。为监督并评估DCLUB中VQA解释的生成质量,我们收集了一个包含1700个聚焦推理问题、附带视觉线索的数据集。评估结果显示,我们的固有可解释系统在聚焦推理问题上的表现比同类黑盒系统提升4.64%,同时在VQA-v2基准上保留了99.43%的性能。
英文摘要:
Recent advances in multimodal large language models (LLMs) have shown extreme effectiveness in visual question answering (VQA). However, the design nature of these end-to-end models prevents them from being interpretable to humans, undermining trust and applicability in critical domains. While post-hoc rationales offer certain insight into understanding model behavior, these explanations are not guaranteed to be faithful to the model. In this paper, we address these shortcomings by introducing an interpretable by design model that factors model decisions into intermediate human-legible explanations, and allows people to easily understand why a model fails or succeeds. We propose the Dynamic Clue Bottleneck Model ( (DCLUB), a method that is designed towards an inherently interpretable VQA system. DCLUB provides an explainable intermediate space before the VQA decision and is faithful from the beginning, while maintaining comparable performance to black-box systems. Given a question, DCLUB first returns a set of visual clues: natural language statements of visually salient evidence from the image, and then generates the output based solely on the visual clues. To supervise and evaluate the generation of VQA explanations within DCLUB, we collect a dataset of 1.7k reasoning-focused questions with visual clues. Evaluations show that our inherently interpretable system can improve 4.64% over a comparable black-box system in reasoning-focused questions while preserving 99.43% of performance on VQA-v2.