arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ViSR-KGC:结合视觉语言模型的视觉子图推理用于多模态知识图谱补全

ViSR-KGC: Visual Subgraph Reasoning with Vision-Language Models for Multimodal Knowledge Graph Completion

Jiafan Li, Mengxue Yang, Jiaqi Zhu, Liang Chang, Ying Li, Hongan Wang

arXiv 2608.05833首次发表:更新:

发表机构

Institute of Software, Chinese Academy of Sciences; University of Chinese Academy of Sciences; CITIC Securities; School of Artificial Intelligence, Beijing Normal University(中国科学院软件研究所; 中国科学院大学; 中信证券; 北京师范大学人工智能学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多模态知识图谱补全的痛点,提出ViSR-KGC方法,结合表示学习、视觉语言模型与预训练常识,通过提取查询感知子图并转换为可视化内容辅助VLM推理缺失实体。

AI 中文摘要

知识图谱补全(KGC)旨在从不完整的图结构中推断缺失的实体或关系,现已发展为多模态知识图谱补全(MMKGC),其中实体与文本、图像等多种模态相关联。传统表示学习方法遵循基于嵌入的范式,在特定关系证据有限时可能表现不佳;同时,基于大语言模型(LLM)的推理方法通常将图结构线性化为文本提示,这会模糊结构拓扑并忽略重要的视觉信息。尽管视觉语言模型(VLMs)在多模态推理方面表现出色,但它们无法直接解释结构化图拓扑,尤其是节点和边带有复杂语义的知识图谱。为弥合这一差距,我们提出ViSR-KGC,一种用于KGC的视觉子图推理方法。它整合了三种互补能力以捕捉语义关联:通过表示学习识别全局拓扑依赖,利用VLMs分析局部多模态证据,以及提供预训练模型中固有的必要常识知识。基于学习到的多模态嵌入,我们的框架首先从MMKG中提取紧凑且查询感知的子图;随后,通过经验选择的布局策略将该子图转换为视觉可解释的图像。在此基础上,将可视化子图、实体图像、文本描述和候选答案组合成统一提示,使VLM能够推断缺失的实体。

英文摘要

Knowledge graph completion (KGC) aims to infer missing entities or relations from incomplete graph structures, and has evolved into multimodal knowledge graph completion (MMKGC), where entities are associated with multiple modalities such as text and images. Traditional representation learning approaches follow the embedding-based paradigm and may struggle when relation-specific evidence is limited. Meanwhile, LLM-based reasoning methods typically linearize graph structures into textual prompts, which obscures structural topology and neglects vital visual information. While vision-language models (VLMs) excel at multimodal reasoning, they cannot natively interpret structured graph topology, particularly when it comes to knowledge graphs where nodes and edges carry complex semantics. To bridge this gap, we propose ViSR-KGC, a visual subgraph reasoning approach for KGC. It integrates three complementary capabilities to capture semantic correlations: identifying global topology dependencies via representation learning, analyzing local multimodal evidence using VLMs, and providing necessary commonsense knowledge inherent in pre-trained models. Based on learned multimodal embeddings, our framework first extracts a compact and query-aware subgraph from the MMKG. Then, this subgraph is transformed into a visually interpretable image using a layout strategy selected through empirical comparison. Finally, the visualized subgraph, entity images, textual descriptions, and candidate answers are combined into a unified prompt, enabling the VLM to infer the missing entity.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑