面向多模态知识图谱的多粒度上下文增强检索增强生成
Multi-Granularity Context-Enhanced RAG over Multimodal Knowledge Graphs
- The Pennsylvania State University(宾夕法尼亚州立大学)
- University of Utah(犹他大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
针对现有多模态知识图谱RAG方法的语义鸿沟问题,本文提出CEMMKG框架,通过多粒度局部与全局上下文增强,有效提升了多模态GraphRAG的性能与适用性。
中文摘要 AI 辅助
检索增强生成(RAG)被广泛用于缓解大语言模型(LLMs)和多模态大语言模型(MLLMs)中的幻觉问题。其中,基于知识图谱(KG)的RAG利用结构化知识为(M)LLMs提供高质量的外部信息。在这些研究的基础上,近期研究探索将多模态知识图谱(MMKGs)作为GraphRAG的知识库,这使得GraphRAG能够整合跨多模态的知识,从而进一步提升其性能。然而,现有的基于MMKG的RAG方法通常遵循一个通用流程:不同模态在融合前大多被独立处理,导致文本上下文在视觉信息提取及后续多模态知识融合过程中仅被有限使用,这造成了图像与文本之间的语义鸿沟,限制了多模态GraphRAG的性能。为解决该问题,本文提出一种用于构建上下文增强型MMKG(CEMMKG)的新型框架,以更好地支持多模态GraphRAG。所提出的CEMMKG在局部和全局两个维度为每张图像补充互补的文本上下文:局部上下文不仅包含图像周围的文本,还纳入与图像语义相关的句子;全局上下文则提供整个段落的摘要。我们进一步为局部上下文引入多粒度设计,使其能够捕捉不同细节层级的语义相关信息。在选定的以视觉为中心的数据集上开展的大量实验验证了,CEMMKG可有效利用上下文信息提升基于MMKG的RAG性能,且其在不同基于MMKG的RAG方法上均表现出有效性,证明了其广泛适用性。
英文摘要
Retrieval-augmented generation (RAG) is widely used to mitigate hallucination issues in large language models (LLMs) and multimodal large language models (MLLMs). In particular, knowledge graph (KG)-based RAG leverages structured knowledge to provide (M)LLMs with high-quality external information. Building on these works, recent studies have explored multimodal knowledge graphs (MMKGs) as knowledge bases for GraphRAG. This enables Graph RAG to integrate knowledge across multiple modalities, thereby further enhancing its performance. However, existing MMKG-based RAG methods generally follow a common pipeline in which different modalities are largely processed independently before being fusion. As a result, textual context is only used to a limited extent during visual information extraction and subsequent multimodal knowledge fusion. This brings a semantic gap between images and text which limits the multimodal GraphRAG performance. To address this issue, we propose a novel framework for constructing a Context-Enhanced MMKG (CEMMKG) to better support multimodal GraphRAG. The proposed CEMMKG enriches each image with complementary textual context at both local and global scopes. Local context goes beyond the surrounding text by incorporating sentences that are semantically related to the image, while global context provides a summary of the entire passage. We further introduce a multi-granularity design for the local context, allowing it to capture semantically relevant information at different levels of detail. Extensive experiments on the selected vision-centric dataset validate that CEMMKG is effective in leveraging contextual information to improve MMKG-based RAG performance. Moreover, its effectiveness across different MMKG-based RAG methods demonstrates its broad applicability.