arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06668cs.LG

迈向统一多模态图基础模型:一种基于桥-路由器-适配器的方法

Towards Unified Multimodal Graph Foundation Model: A Bridge-Router-Adapter Based Approach

  • Beijing Institute of Technology(北京理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Sirui Zhang, Yubing Zhou, Xunkai Li, Zekai Chen, Shumeng Li, Wang Luo, Yinlin Zhu, Yujin Gao, Rong-Hua Li

中文总结 AI 辅助

针对多模态图基础模型中上下文纠缠与模态路由局限,提出基于桥-路由器-适配器的统一模型BRAIN,通过范围条件化桥、分层路由器和残差适配器,在多个数据集上显著提升性能。

中文摘要 AI 辅助

多模态图将不同模态(如文本和图像)的节点属性与关系结构耦合,从而能够联合建模拓扑结构和跨模态属性。多模态图基础模型旨在从这类数据中学习统一的表示,并使其能够跨不同的图领域和下游任务进行迁移。然而,现有方法存在两个根本性局限。(1)跨范围上下文纠缠:它们将特定范围的图上下文合并为统一表示,在多模态构建过程中模糊了它们之间的区别。(2)忽略范围的多模态路由:它们在固定的图范围内路由模态,忽略了模态相关性随邻域范围的变化。为应对这些挑战,我们提出了BRAIN,一个统一模型,其聚焦于结合邻域范围与模态组成的图上下文。BRAIN包含一个范围条件化的桥(Bridge),将跨越从局部到全局邻域范围的结构信息与不同的模态组成相结合;一个分层路由器(Router),估计范围与任务之间的相关性,并在每个范围内分别选择组成,使模态效用随图范围变化;以及一个轻量级残差适配器(Adapter),进一步将路由后的嵌入专门化以用于下游预测。BRAIN通过多图预训练后进行任务特定微调来训练。在九个数据集和四个任务族上的实验证明了其广泛的有效性,相对于最强基线,节点分类和链接预测性能最高提升了4.73%,同时在四个图到文本和两个图到图像指标上实现了平均14.72%的相对提升。

英文摘要

Multimodal graphs couple node attributes in different modalities, such as text and images, with relational structure, enabling topological structure and cross-modality attributes to be modeled jointly. Multimodal graph foundation models seek unified representations from such data that transfer across different graph domains and downstream tasks. However, existing methods exhibit two fundamental limitations. (1) Cross-Scope Context Entanglement. They merge scope-specific graph contexts into a unified representation, obscuring their distinctions during multimodal construction. (2) Scope-Ignorant Modality Routing. They route modalities within a fixed graph scope, overlooking how modality relevance varies across neighborhood ranges. To address these challenges, we propose BRAIN, a unified model that focuses on graph context that combines neighborhood scope with modality composition. BRAIN comprises a scope-conditioned Bridge that combines structural information spanning local-to-global neighborhood scopes with different modality compositions; a hierarchical Router that estimates the relevance between the scope and the task, and selects compositions separately within each scope, allowing modality utility to vary with graph range; and a lightweight residual Adapter that further specializes the routed embedding for downstream prediction. BRAIN is trained through multi-graph pretraining followed by task-specific adaptation. Experiments across nine datasets and four task families demonstrate its broad effectiveness, improving node-classification and link-prediction performance by up to 4.73% relative to the strongest baseline, while achieving an average relative improvement of 14.72% across four graph-to-text and two graph-to-image metrics.

↑