发表机构
School of Computer Science and Technology, College of Intelligence and Computing, Tianjin University; Data Science Program, Columbian College of Arts and Sciences, The George Washington University(天津大学智能与计算学部计算机科学与技术学院; 乔治华盛顿大学文理学院哥伦比亚艺术与科学学院数据科学项目)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究多模态图零样本迁移问题,提出CHARM模型,通过分层上下文建模,用图上下文替换原始节点,减少对目标域监督或适配的依赖,经实验验证该模型在零样本多模态图任务上有改进。
AI 中文摘要
图基础模型已成为跨图域和任务迁移知识的有前景的范式。现实世界的图将节点与文本、图像等模态相关联,使得多模态图对于表示复杂实体和关系至关重要。由于收集标签和为每个新图域调整模型成本高且往往不可行,因此激发了零样本迁移的需求。然而,多模态图上的零样本迁移仍未得到充分探索。现有基于GNN的图基础模型通常需要下游适配,而基于LLM的图方法主要处理单模态图或单个域内的任务。本文提出了CHARM,一种用于零样本迁移的具有分层上下文建模的多模态图基础模型。CHARM用捕获多模态语义和跨模态关系的分层图上下文替换孤立的原始节点。这些上下文将特定域的节点模式映射到共享的高级概念,减少对目标域监督或适配的依赖。实验表明在零样本多模态图任务上有持续改进。
英文摘要
Graph foundation models (GFMs) have emerged as a promising paradigm for transferring knowledge across graph domains and tasks. Real-world graphs associate nodes with text, images, and other modalities, making multimodal graphs essential for representing complex entities and relations. Moreover, collecting labels and adapting models for every new graph domain is costly and often infeasible, motivating zero-shot transfer. Unfortunately, zero-shot transfer on multimodal graphs remains underexplored. Existing GNN-based graph foundation models typically require downstream adaptation, whereas LLM-based graph methods mainly address unimodal graphs or tasks within a single domain. This setting presents two key challenges. First, models must generalize knowledge from individual modalities while capturing transferable cross-modal relations. Second, without target-domain fine-tuning, node representations remain entangled with domain-specific structures and modality-specific characteristics, obscuring shared concepts in unseen domains. To address these challenges, we propose CHARM, a multimodal graph foundation model with hierarchical context modeling for zero-shot transfer. CHARM replaces isolated raw nodes with hierarchical graph contexts that capture multimodal semantics and cross-modal relations. These contexts map domain-specific node patterns to shared high-level concepts, reducing reliance on target-domain supervision or adaptation. A modality-aware graph context encoder integrates multimodal information with graph structure and converts the resulting representations into graph tokens for a large language model . Experiments show consistent improvements on zero-shot multimodal graph tasks.