arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ReCoG:面向多模态图学习的互惠协同演化

ReCoG: Reciprocal Co-Evolution for Multimodal Graph Learning

Rui Xue, Tianfu Wu

arXiv 2608.22786首次发表:更新:

发表机构

North Carolina State University(北卡罗来纳州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ReCoG是将图结构学习与多模态表示学习紧密耦合的新范式,在节点分类和链接预测基准中优于多模态图结构学习基线,证明了结构与语义协同演化的重要性。

AI 中文摘要

多模态图学习需要对图结构与异构节点属性进行联合训练,但现有方法大多将这两个过程解耦:现有的多模态图神经网络(GNN)聚焦于在共享嵌入空间中对齐模态,同时在固定或弱适配的图结构上运行;而图结构学习方法则从单模态节点表示中推断拓扑,未考虑多模态交互。这种分离从根本上限制了GNN在多模态场景中捕捉语义上有意义关系的能力,在该场景中观测到的边往往存在噪声、不完整或与底层语义不匹配。我们提出ReCoG(面向多模态图学习的互惠协同演化,Reciprocal Co-Evolution for Multimodal Graph Learning),这是一种新的学习范式,通过端到端的互惠交互将图结构学习与多模态表示学习紧密耦合。具体而言,ReCoG集成了:(i)一个多模态图精化器,利用跨模态语义证据推断并修正边;(ii)一个耦合跨模态消息传递机制,在精化后的图上执行模态内与跨模态的联合传播。这种统一设计比解耦或两阶段公式具有更强的表达能力,并允许拓扑与表示学习之间的动态交互。在节点分类和链接预测的各类基准测试中,ReCoG始终优于强大的多模态图结构学习基线,包括图基础模型。我们的结果表明,结构与语义的互惠协同演化对有效的多模态图学习至关重要,挑战了拓扑与表示学习之间普遍存在的分离。

英文摘要

Multimodal graph learning requires jointly training over graph structure and heterogeneous node attributes, yet existing methods largely decouple these processes: prior multimodal graph neural networks (GNNs) focus on aligning modalities in a shared embedding space while operating on fixed or weakly adapted graph structures, and graph structure learning approaches infer topology from unimodal node representations without accounting for multimodal interactions. This separation fundamentally limits the ability of GNNs to capture semantically meaningful relationships in multimodal settings, where observed edges are often noisy, incomplete, or misaligned with underlying semantics. We propose ReCoG (Reciprocal Co-Evolution for Multimodal Graph Learning), a new learning paradigm that tightly couples graph structure learning and multimodal representation learning through end-to-end reciprocal interaction. Concretely, ReCoG integrates (i) a multimodal graph refiner that infers and corrects edges using cross-modal semantic evidence, and (ii) a coupled cross-modal message passing mechanism that performs joint intra- and inter-modality propagation over the refined graph. This unified design yields greater expressiveness than decoupled or two-stage formulations and allows dynamic interaction between topology and representation learning. Across diverse benchmarks for node classification and link prediction, ReCoG consistently outperforms strong multimodal graph structure learning baselines, including graph foundation models. Our results demonstrate that reciprocal co-evolution of structure and semantics is important for effective multimodal graph learning, challenging the prevailing separation between topology and representation learning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑