arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.21155cs.CVcs.AI

CRAG-MM诊断:实现对知识密集型视觉问答的逐阶段分析

CRAG-MM-Diagnostics: Enabling Stage-Wise Analysis of Knowledge-Intensive VQA

Hanseok Oh, Parishad BehnamGhader, Benno Krojer, Hyunji Lee, Paul Liang, Siva Reddy, Verna Dankers

首次发表
浏览论文内容

中文总结 AI 辅助

研究知识密集型视觉问答流程,引入CRAG-MM-Diagnostics诊断基准,通过逐阶段数据注释分离子问题,评估VLM并进行细粒度分析,指出知识检索和推理是主要瓶颈,还提出改进流程提升准确率。

中文摘要 AI 辅助

知识密集型视觉问答(KI-VQA)基准通过要求提供图像之外的外部信息来评估视觉语言模型(VLM)作为多模态知识助手。KI-VQA涉及多个子问题,如指代表达理解、视觉定位、目标识别、知识检索和推理等,但现有基准通常只报告最终任务的准确率,掩盖了失败发生的位置。为了分析完整的KI-VQA流程,我们引入了CRAG-MM-Diagnostics,这是一个具有逐阶段数据注释的诊断基准,可分离出基于语言的视觉定位、目标识别以及知识检索和推理。我们评估了完全参数化和检索增强的VLM,并使用新收集的元数据(如目标感兴趣区域、实体名称和视觉复杂度分数)进行细粒度分析。结果表明知识检索和推理是主要瓶颈,同时也凸显了KI-VQA流程其他部分的问题,并利用这些发现提出了一个基于视觉定位模块的双模态RAG流程,提高了GPT-5和Qwen的准确率。

英文摘要

Knowledge-Intensive Visual Question Answering (KI-VQA) benchmarks evaluate Vision-Language Models (VLMs) as multimodal knowledge assistants by requiring external information beyond a provided image to answer questions. KI-VQA involves multiple sub-problems -referring expression understanding, visual grounding, object recognition, knowledge retrieval, and reasoning-yet existing benchmarks typically report only end-task accuracy, obscuring where failures arise. To analyze the full KI-VQA pipeline, we introduce CRAG-MM-Diagnostics, a diagnostic benchmark with stage-wise data annotations that isolate 1) language-based visual grounding, 2) object identification, and 3) knowledge retrieval and reasoning. We evaluate fully parametric and retrieval-augmented VLMs, providing fine-grained analyses using newly collected metadata, such as target ROIs, entity names, and visual complexity scores. Our results point to knowledge retrieval and reasoning as the primary bottleneck, but also highlight issues in the other parts of the KI-VQA pipeline, such as the fact that VLMs struggle with target object identification or that image retrievers struggle to integrate textual cues. These findings expose fundamental limitations in current KI-VQA systems and motivate stage-aware evaluation. We, lastly, leverage these findings to propose a grounded bimodal RAG pipeline that integrates a visual grounding module to crop targets before image retrieval, boosting GPT-5 and Qwen's respective accuracies by 13.3 and 8.5 percentage points.

发表机构

  • New York University(纽约大学)
  • McGill University(麦吉尔大学)
  • Mila - Quebec AI Institute(米拉-魁北克人工智能研究所)
  • UNC Chapel Hill(北卡罗来纳大学教堂山分校)
  • MIT(麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑