arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2606.26458cs.AI

MKG-RAG-Bench:多模态知识图谱增强生成中的检索基准

MKG-RAG-Bench: Benchmarking Retrieval in Multimodal Knowledge Graph-Augmented Generation

  • The Pennsylvania State University(宾夕法尼亚州立大学)
  • Michigan State University(密歇根州立大学)
  • Dalian University of Technology(大连理工大学)
  • Stony Brook University(石溪大学)

机构由 AI 辅助整理,请以论文原文为准。

Xiaochen Wang, Bao Hoang, Han Liu, Ting Wang, Fenglong Ma

AI总结:

提出MKG-RAG-Bench基准,通过构建跨领域多模态知识图谱和问答数据集,系统评估多模态知识图谱增强生成中的检索性能,揭示检索质量对生成结果的决定性作用。

AI中文摘要:

基于知识图谱的检索增强生成(RAG)已成为一种有前景的方法,用于支撑大型语言模型,但现有基准大多忽视了多模态知识图谱RAG(MKG-RAG)中检索的挑战。在实践中,检索是一个关键瓶颈:多模态知识是异质的,难以跨模态对齐,并且通常由为无结构语料库设计的检索器服务不佳。为了解决这一差距,我们引入了MKG-RAG-Bench,这是一个明确设计用于评估MKG-RAG中检索的跨领域基准。MKG-RAG-Bench由两个覆盖通用和医学领域的多模态知识图谱构建而成,并包含精心对齐的问答数据集,支持对检索和下游生成进行受控评估。该基准使用基于LLM的策划流程构建,该流程过滤低效用知识,生成具有精确监督的结构化查询,并系统地覆盖多种模态配置。通过在代表性检索器系列和模态设置上的广泛实验,我们表明有效的多模态检索仍然具有挑战性,但对端到端MKG-RAG性能至关重要,并且检索质量强烈决定生成结果。通过将检索作为首要评估目标,MKG-RAG-Bench为诊断当前限制和推进多模态知识图谱RAG系统提供了原则性基础。

英文摘要:

Retrieval-augmented generation (RAG) over knowledge graphs has emerged as a promising approach for grounding large language models, yet existing benchmarks largely overlook the challenges of retrieval in multimodal knowledge graph RAG (MKG-RAG). In practice, retrieval is a critical bottleneck: multimodal knowledge is heterogeneous, difficult to align across modalities, and often poorly served by retrievers designed for unstructured corpora. To address this gap, we introduce MKG-RAG-Bench, a cross-domain benchmark explicitly designed to evaluate retrieval in MKG-RAG. MKG-RAG-Bench is constructed from two multimodal knowledge graphs spanning general and medical domains, and includes carefully aligned question-answering datasets that support controlled evaluation of both retrieval and downstream generation. The benchmark is built using an LLM-based curation pipeline that filters low-utility knowledge, generates structurally grounded queries with exact supervision, and systematically covers diverse modality configurations. Through extensive experiments across representative retriever families and modality settings, we show that effective multimodal retrieval remains challenging yet crucial for end-to-end MKG-RAG performance, and that retrieval quality strongly determines generation outcomes. By isolating retrieval as a first-class evaluation target, MKG-RAG-Bench provides a principled foundation for diagnosing current limitations and advancing multimodal knowledge graph RAG systems.

补充信息

↑