CrossModalQA:面向多模态检索增强生成的多模态多跳基准
CrossModalQA: A Cross-modal and Multi-hop Benchmark for Multimodal Retrieval-augmented Generation
浏览论文内容
中文总结 AI 辅助
提出开放域多模态多跳基准CrossModalQA,含1,863个问答对,覆盖五种跨模态推理路径,实验表明完整跨模态检索比生成器规模更关键,多图像推理是主要瓶颈。
中文摘要 AI 辅助
尽管多模态大语言模型(MLLMs)能力强大,但其参数化知识仍不完整且难以更新,这促使多模态检索增强生成(RAG)利用外部文本和图像来支撑响应。然而,现有基准存在两大局限:(i)它们通常侧重于单跳检索或在少量给定上下文上的推理,而非开放域证据发现;(ii)它们对跨模态推理路径的覆盖零散,使得复杂的多跳和多图像推理未得到充分探索。本文提出CrossModalQA,一个用于评估异构语料库上多模态检索与推理的开放域基准。CrossModalQA包含1,863个问答对,由4,987篇维基百科文章和4,431张维基共享资源图像构建而成。它涵盖五种互补的推理路径:视觉到文本、文本到视觉、视觉到文本到视觉、多图像交集和图像集推理。每个问题都需要检索并组合分布式的文本和视觉证据,平均推理深度为3.50跳。我们通过多模态知识图谱引导的子图采样构建该基准,并应用基于规则的一致性检查和LLM验证,以确保多模态依赖性和可追溯的证据。大量实验表明,现有的多模态RAG系统难以恢复完整的证据链,且当不完整检索引入干扰性上下文时,其表现可能不如闭卷模型。进一步分析揭示,完整的跨模态检索对答案准确性的贡献大于生成器规模的扩大,而多图像检索与推理仍是限制端到端性能的主要瓶颈。
英文摘要
Despite the strong capabilities of multimodal large language models (MLLMs), their parametric knowledge remains incomplete and difficult to update, motivating multimodal retrieval-augmented generation (RAG) to ground responses in external text and images. However, existing benchmarks face two major limitations: (i) they typically emphasize single-hop retrieval or reasoning over a small set of provided contexts rather than open-domain evidence discovery; and (ii) they provide fragmented coverage of cross-modal reasoning paths, leaving complex multi-hop and multi-image reasoning underexplored. In this paper, we introduce CrossModalQA, an open-domain benchmark for evaluating multimodal retrieval and reasoning over heterogeneous corpora. CrossModalQA contains 1,863 question-answer pairs constructed from 4,987 Wikipedia articles and 4,431 Wikimedia Commons images. It covers five complementary reasoning paths: vision-to-text, text-to-vision, vision-to-text-to-vision, multi-image intersection, and image-set reasoning. Every question requires retrieving and composing distributed textual and visual evidence, with an average reasoning depth of 3.50 hops. We construct the benchmark through multimodal knowledge graph-guided subgraph sampling and apply rule-based consistency checking and LLM verification to ensure multimodal dependence and traceable evidence. Extensive experiments demonstrate that existing multimodal RAG systems struggle to recover complete evidence chains and can underperform closed-book models when incomplete retrieval introduces distracting context. Further analysis reveals that complete cross-modal retrieval contributes more to answer accuracy than generator scaling, while multi-image retrieval and reasoning remain the primary bottlenecks limiting end-to-end performance.
发表机构
- The Hong Kong Polytechnic University(香港理工大学)
- Jilin University(吉林大学)
机构由 AI 辅助整理,请以论文原文为准。