arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

信号还是噪声?多模态GraphRAG中的模态贡献与协作

Signal or Noise? Modality Contribution and Cooperation in Multimodal GraphRAG

Antonios Georgakopoulos, Paul Groth, Lise Stork

arXiv 2609.35304首次发表:更新:

发表机构

University of Amsterdam(阿姆斯特丹大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究以DocVQA为测试平台,通过模态感知的多模态GraphRAG框架分析模态贡献与协作,发现表格文本贡献最强、模态组合多冗余,提出应按任务选择性检索模态以优化系统。

AI 中文摘要

多模态知识图谱(KGs)将文本、图表、表格及其他模态的信息整合为统一的结构化表示,有望通过更丰富的证据实现更优推理。在基于此类图谱构建的GraphRAG系统中,人们普遍认为推理时从更多模态检索证据能提升下游任务性能。然而,冗余或重叠的多模态证据可能会干扰语言模型的问答(QA)过程,且各模态在不同问题、模型和任务中的贡献是否均等仍缺乏深入研究。\n本研究以文档视觉问答(DocVQA)为测试平台,探究模态感知检索对多模态GraphRAG流水线下游推理的影响。我们对现有基于KG的问答框架进行扩展,使其具备模态感知能力,利用图结构追踪各事实对应的支撑模态,并在边级别选择性过滤证据。这使我们能够验证推理时提供所有可用多模态证据是否对问答任务有益,并评估不同模态在问题、任务和模型特征维度下的贡献与协作模式。\n通过在前沿多模态GraphRAG流水线、5个多模态大语言模型(LLMs)及2个DocVQA基准上开展对照分析,我们发现表格和文本的贡献最为突出,模态组合往往产生冗余而非协同效应,涉及文本信息的模态对尤其如此。正向协作主要出现在非文本模态之间,且取决于问题意图和任务类型。研究结果表明,设计更高效的GraphRAG系统应采用选择性的模态感知检索策略,根据下游任务过滤模态而非统一检索所有模态。

英文摘要

Multimodal knowledge graphs (KGs) integrate information from text, figures, tables, and other modalities into a unified structured representation, with the promise that richer evidence enables better inference. In GraphRAG systems built over such graphs, it is commonly assumed that retrieving evidence from more modalities at inference time improves downstream performance. Yet, redundant or overlapping multimodal evidence may distract language models in question answering (QA), and whether each modality contributes equally across questions, models, and tasks remains poorly understood. In this work, we study how modality-aware retrieval affects downstream inference in a multimodal GraphRAG pipeline, using document visual question answering (DocVQA) as a testbed. We extend an existing KG-based QA framework to be modality-aware, leveraging the graph structure to track which modality supports which facts and to selectively filter evidence at the edge level. This enables us to investigate whether providing all available multimodal evidence at inference time benefits QA, and to evaluate the contribution and cooperation of modalities across question, task, and model characteristics. Through a controlled analysis within a state-of-the-art multimodal GraphRAG pipeline, five multimodal LLMs and two DocVQA benchmarks, we find that tables and text provide the strongest contributions, and that combining modalities frequently produces redundancy rather than synergy, particularly for pairs involving textual information. Positive cooperation appears mainly between non-text modalities and depends on question intent and task type. Our findings argue for selective, modality-aware retrieval in the design of more effective GraphRAG systems, where modalities are filtered according to the downstream task rather than retrieved uniformly.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑