arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MGDT:具有关系自适应专家混合的MLLM引导扩散变压器用于多模态知识图谱补全

MGDT: MLLM-Guided Diffusion Transformer with Relation-Adaptive Mixture-of-Experts for Multimodal Knowledge Graph Completion

Xu Hou, Meiyu Liang, Wei Huang, Yawen Li, Zhe Xue, Wu Liu, Guanhua Ye, Lei Shi, Kangkang Lu

arXiv 2607.15592首次发表:更新:

发表机构

Beijing University of Posts and Telecommunications; Zhejiang University; University of Science and Technology of China; Communication University of China(北京邮电大学; 浙江大学; 中国科学技术大学; 中国传媒大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究多模态知识图谱补全问题,提出MGDT框架,先通过RASR-MoE模块选路径、抑干扰,再用MLLM对齐表示,最后KGDT去噪生成,实验证明该框架在三个基准数据集上性能优于基线。

AI 中文摘要

多模态知识图谱补全(MKGC)需要从结构、文本和视觉线索中推断缺失实体。现有的基于扩散的MKGC方法通常直接对原始多模态特征去噪。这种设计使去噪器同时执行依赖关系的线索选择、跨模态语义对齐和结构感知实体生成,给扩散带来噪声和语义不一致条件,导致补全性能次优。为解决此限制,我们提出MGDT,一种基于先对齐后扩散范式构建的新型MKGC框架。MGDT首先使用关系自适应语义路由专家混合(RASR-MoE)模块选择与关系相关的多模态语义转换路径并抑制无关模态干扰。然后使用冻结的多模态大语言模型(MLLM)作为语义锚将路由后的多模态表示对齐到统一潜在空间并减少跨模态语义异质性。最后,知识图谱扩散变压器(KGDT)在对齐空间中执行图条件去噪生成以产生缺失实体表示。在三个基准数据集上的实验表明,MGDT始终优于强大的基线。

英文摘要

Multimodal Knowledge Graph Completion (MKGC) requires inferring missing entities from structural, textual, and visual cues. Existing diffusion-based MKGC methods usually denoise directly on raw multimodal features. Such a design forces the denoiser to simultaneously perform relation-dependent cue selection, cross-modal semantic alignment, and structure-aware entity generation, which introduces noisy and semantically inconsistent conditions for diffusion and consequently leads to suboptimal completion performance. To address this limitation, we propose MGDT: MLLM-Guided Diffusion Transformer with Relation-Adaptive Mixture-of-Experts (MGDT), a novel MKGC framework built on an align-then-diffuse paradigm. MGDT first employs a Relation-Adaptive Semantic Routing Mixture-of-Experts (RASR-MoE) module to select relation-relevant multimodal semantic transformation paths and suppress irrelevant modality interference. MGDT then uses a frozen Multimodal Large Language Model (MLLM) as a semantic anchor to align the routed multimodal representations into a unified latent space and reduce cross-modal semantic heterogeneity. Finally, a Knowledge Graph Diffusion Transformer (KGDT) performs graph-conditioned denoising generation in the aligned space to produce the missing entity representation. Experiments on three benchmark datasets show that MGDT consistently outperforms strong baselines.

Comments8pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑