arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

迈向全模态图基础模型:一种拓扑驱动的绑定方法

Toward Omni Multimodal Graph Foundation Model: A Topology-Driven Binding Approach

Xunkai Li, Chenxi Wan, Yinlin Zhu, Wang Luo, Hongchao Qin, Rong-Hua Li, Guoren Wang

arXiv 2610.02881首次发表:更新:

发表机构

Beijing Institute of Technology; Sun Yat-sen University(北京理工大学; 中山大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多模态图基础模型在属性不完整及拓扑利用不足的问题,提出拓扑驱动的GraphBind方法,将模态信息绑定至统一空间,在判别与生成任务上相对最强基线提升达28.1%。

AI 中文摘要

多模态图基础模型(MGFMs)旨在从具有异构节点模态的大规模图中学习可泛化的表示。然而,现实世界中的多模态属性图(MAGs)往往包含不完整的节点属性,限制了可用预训练语料库的规模和多样性。此外,现有的MGFMs主要将图拓扑作为结构上下文纳入,忽视了其在引导多模态绑定和塑造统一表示空间方面的作用。为应对这些挑战,我们提出了GraphBind,一种拓扑驱动的方法,利用图拓扑将丰富的模态信息绑定到统一的共享空间中。GraphBind的动机源于图拓扑的稳定性,它为多模态绑定提供了结构参考和互补的语义信息。具体而言,GraphBind利用拓扑将自身语义和可靠的邻域语义组织到一个集成结构与语义的全局共享空间中,并通过轻量级接口使该空间适应判别性和生成性任务。针对11个代表性基线的广泛实验表明,GraphBind在判别性和生成性任务上均取得了领先性能,相对于最强基线实现了高达28.1%的相对提升。

英文摘要

Multimodal graph foundation models (MGFMs) seek to learn generalizable representations from large-scale graphs with heterogeneous node modalities. However, real-world Multimodal-Attributed Graphs (MAGs) often contain incomplete node attributes, limiting the scale and diversity of available pretraining corpora. Besides, existing MGFMs primarily incorporate graph topology as structural context, overlooking its role in guiding multimodal binding and shaping a unified representation space. To address these challenges, we propose GraphBind, a topology-driven approach that uses graph topology to bind rich modality information into a unified shared space. GraphBind is motivated by the stability of graph topology, which provides structural references and complementary semantic information for multimodal binding. Concretely, GraphBind uses topology to organize self semantics and reliable neighborhood semantics into a global shared space that integrates structure and semantics, and adapts this space to discriminative and generative tasks through lightweight interfaces. Extensive experiments against 11 representative baselines demonstrate that GraphBind achieves leading performance on both discriminative and generative tasks, with relative improvements of up to 28.1% over the strongest baseline.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑