LION:面向多模态属性图学习的克利福德神经范式
LION: A Clifford Neural Paradigm for Multimodal-Attributed Graph Learning
- Beijing Institute of Technology(北京理工大学)
- Ant Group(蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对多模态属性图学习中模态对齐忽略上下文、模态融合缺乏适应性的问题,提出基于克利福德代数和解耦图神经范式的LION方法,在9个文本-图像多模态属性图数据集上的多类下游任务中显著优于现有最优基线。
AI中文摘要:
近期,多模态领域的快速发展推动了图机器学习(graph ML)中以数据为中心的范式转变,从文本属性图过渡到多模态属性图。这一进展显著提升了数据表示能力,拓展了图下游任务的范围,例如面向模态的任务,从而增强了图机器学习的实用价值。尽管前景广阔,但当前的神经范式存在局限性:(1)模态对齐中忽略上下文:大多数现有方法采用受拓扑约束或模态特定的算子作为对齐器,不可避免地忽略图上下文并抑制模态交互,导致对齐效果欠佳;(2)模态融合中缺乏适应性:大多数现有方法仅针对双模态图进行简单适配,未能在融合过程中充分利用带有拓扑先验的对齐令牌,导致泛化能力和性能较差。为解决上述问题,我们提出基于克利福德代数(Clifford algebra)和解耦图神经范式(即先传播后聚合)的LION(克利福德神经范式),以在多模态属性图中实现先对齐后融合。具体而言,我们首先构建基于克利福德代数的模态感知几何流形,这种几何诱导的高阶图传播能有效实现模态交互,促进模态对齐。接着,基于对齐令牌的拓扑感知克利福德分量,我们提出自适应全息聚合模块,该模块结合分量能量和传播尺度信息与可学习参数,以改进模态融合。在9个文本-图像多模态属性图(MAG)数据集上开展的大量实验表明,LION在3项图下游任务和3项模态下游任务中均显著优于现有最优(SOTA)基线。
英文摘要:
Recently, the rapid advancement of multimodal domains has driven a data-centric paradigm shift in graph ML, transitioning from text-attributed to multimodal-attributed graphs. This advancement significantly enhances data representation and expands the scope of graph downstream tasks, such as modality-oriented tasks, thereby improving the practical utility of graph ML. Despite its promise, limitations exist in the current neural paradigms:(1) Neglect Context in Modality Alignment: Most existing methods adopt topology-constrained or modality-specific operators as tokenizers.These aligners inevitably neglect graph context and inhibit modality interaction, resulting in suboptimal alignment.(2) Lack of Adaptation in Modality Fusion: Most existing methods are simple adaptations for 2-modality graphs and fail to adequately exploit aligned tokens equipped with topology priors during fusion, leading to poor generalizability and performance degradation.To address the above issues, we propose LION (c\underline{LI}ff\underline{O}rd \underline{N}eural paradigm) based on the Clifford algebra and decoupled graph neural paradigm (i.e., propagation-then-aggregation) to implement alignment-then-fusion in multimodal-attributed graphs. Specifically, we first construct a modality-aware geometric manifold grounded in Clifford algebra.This geometric-induced high-order graph propagation efficiently achieves modality interaction, facilitating modality alignment.Then, based on the topology-aware Clifford components of aligned tokens, we propose adaptive holographic aggregation. This module integrates component-wise energy and propagation-scale information with learnable parameters to improve modality fusion. Extensive experiments on 9 text-image MAG datasets demonstrate that LION significantly outperforms SOTA baselines across 3 graph and 3 modality downstream tasks.