发表机构
MBZUAI The University of Melbourne(MBZUAI墨尔本大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究视觉语言模型在回答需额外知识问题时的错误,提出统一框架分离失败模式,探讨预生成信号能否预测错误源,发现可在解码前预测,能据此进行针对性干预。
AI 中文摘要
视觉语言模型在高质量图像的视觉问答上表现良好,但在回答需要超出清晰直接可见知识的问题时存在困难。本文提出统一框架来解开这些失败模式,并研究预生成信号是否能预测这些失败源。发现在一系列数据集和模型家族中,VLM错误存在一致模式,且可在解码前预测失败源,实现针对性干预。
英文摘要
Vision-language models (VLMs) can recognize entities in clear images yet still fail when answering questions that require factual knowledge beyond what is directly observable. Prior work has either examined individual failure modes in isolation or treated incorrect answers as monolithic, binary failures. We propose a tree-structured framework that organizes failures in knowledge-intensive visual question answering into model-specific operational outcomes. Across two datasets and four VLMs, we observe consistent distributions of operational outcomes: some failures occur before entity recognition, while others persist after the relevant entity is recognized. Visual token representations are most informative for recognition-related decisions. Prompt hidden states predict answer success more effectively, although factual-access attribution remains difficult and exhibits only a weak signal. These pre-generation signals support attribution-guided routing to targeted interventions, including image repair, entity support, question rewriting, and factual evidence.