发表机构
New York University Shanghai; New York University(纽约大学上海分校; 纽约大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究利用视觉语言模型解决属性图模态异质性问题,提出OMG-VLM框架,以预训练VLM为共享主干,引入结构感知图适配器,在节点分类和链接预测等任务中表现出色,优于现有基线且泛化能力强。
AI 中文摘要
视觉语言模型(VLM)为文本和视觉信息提供了统一的表示空间,但其作为图结构数据通用主干的潜力尚未得到充分探索。属性图存在显著的模态异质性,现有图学习方法通常针对固定模态模式设计,需要为不同设置单独建模,限制了可扩展性和跨图泛化。为此提出OMG-VLM框架,利用预训练VLM作为共享主干,引入结构感知图适配器整合邻域信息并与VLM原生嵌入空间兼容。实验表明,OMG-VLM在属性图学习任务上优于现有基线,且具有强泛化能力。
英文摘要
Vision-language models (VLMs) provide a unified representation space for textual and visual information, yet their potential as general-purpose backbones for graph-structured data remains largely unexplored. In practice, attributed graphs exhibit substantial modality heterogeneity: some graphs contain only textual node attributes, others only visual attributes, while still others provide both. Existing graph learning approaches are typically designed for fixed modality schemas, requiring separate models for different settings and limiting scalability and cross-graph generalization. To bridge this gap, we present OMG-VLM (One Model, Many Graphs with Vision-Language Models), a unified framework for learning over attributed graphs across heterogeneous modality schemas. OMG-VLM leverages a pretrained VLM as a shared backbone and introduces structure-aware graph adapters that integrate neighborhood information while remaining compatible with the VLM's native embedding space. This design enables effective learning over text-attributed, image-attributed, and multimodal-attributed graphs within a single model. Extensive experiments across diverse domains show that OMG-VLM consistently outperforms state-of-the-art GNN- and LLM-based baselines on attributed graph learning tasks such as node classification and link prediction, while exhibiting strong generalization to unseen graphs and varying modality schemas. The source code is available at https://github.com/Jo-eyang/OMG-VLM.
CommentsEMNLP 2026 Main Conference