arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

是看不见还是不知道?归因视觉语言模型中的错误

Failing to See or Failing to Know? Attributing Errors in Vision-Language Models

Khang Nhat Hoang Vo, Artem Vazhentsev, Artem Shelmanov, Timothy Baldwin, Yova Kementchedjhieva

arXiv 2607.04683首次发表:更新:

发表机构

MBZUAI The University of Melbourne(MBZUAI墨尔本大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究视觉语言模型在回答需额外知识问题时的错误,提出统一框架分离失败模式,探讨预生成信号能否预测错误源,发现可在解码前预测,能据此进行针对性干预。

AI 中文摘要

视觉语言模型在高质量图像的视觉问答上表现良好,但在回答需要超出清晰直接可见知识的问题时存在困难。本文提出统一框架来解开这些失败模式,并研究预生成信号是否能预测这些失败源。发现在一系列数据集和模型家族中,VLM错误存在一致模式,且可在解码前预测失败源,实现针对性干预。

英文摘要

Vision-language models (VLMs) can recognize entities in clear images yet still fail when answering questions that require factual knowledge beyond what is directly observable. Prior work has either examined individual failure modes in isolation or treated incorrect answers as monolithic, binary failures. We propose a tree-structured framework that organizes failures in knowledge-intensive visual question answering into model-specific operational outcomes. Across two datasets and four VLMs, we observe consistent distributions of operational outcomes: some failures occur before entity recognition, while others persist after the relevant entity is recognized. Visual token representations are most informative for recognition-related decisions. Prompt hidden states predict answer success more effectively, although factual-access attribution remains difficult and exhibits only a weak signal. These pre-generation signals support attribution-guided routing to targeted interventions, including image repair, entity support, question rewriting, and factual evidence.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑