ICE:面向多模态图基础模型的任务对齐Clifford潜空间
ICE: Task-Aligned Clifford Latent Fields for Multimodal Graph Foundation Models
- Beijing Institute of Technology(北京理工大学)
- Shandong University(山东大学)
- Sun Yat-sen University(中山大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出ICE,一种基于Clifford潜空间的多模态图基础模型,通过任务对齐的几何表示保留实体语义并构建交互状态,在30个监督和小样本任务中均排名第一。
AI中文摘要:
多模态属性图连接实体、视觉内容、语言和观测到的关系。在此类图上学习一个基础模型,不仅仅需要将每个节点压缩为融合的欧几里得向量。表示必须保留实体语义,从图邻域构建交互状态,并将该状态暴露给具有不同几何结构的预测单元。我们的实证研究表明,这些要求为何密不可分。更高级别的通道在基础图之间恢复成对关系,专门的查询揭示通用读出隐藏的信息,而刚性的刀片隔离消除了跨级别的容量。因此,我们引入ICE(交互感知Clifford编码器),一种基于节点索引的Clifford潜空间构建的多模态图基础模型。拓扑、文本和图像进入显式的Cl(3)地址。边感知几何积将这些方向转换为观测邻域上的标量、双向量和三向量关系。受保护的Grade-1路径保留实体语义,而完整的级别和深度库仍可用于新的节点和链接头。我们建立了精确的跨级别可达性、节点排列等变性和围绕语义分数的任务残差界限。实验涵盖一个共享基础模型,跨越11个图、6个节点分类数据集、3个链接预测数据集和匹配的小样本任务。ICE在所有30个报告的监督和小样本比较中排名第一。核心移除降低了每个任务摘要,机制控制将收益与高阶传输、保留的多深度结构、语义保护和直接字段访问联系起来。
英文摘要:
Multimodal attributed graphs connect entities, visual content, language, and observed relations. Learning one foundation across such graphs requires more than compressing each node into a fused Euclidean vector. The representation must preserve entity semantics, construct interaction state from graph neighborhoods, and expose that state to prediction units with different geometry. Our empirical study shows why these requirements are inseparable. Higher-grade channels recover pair relations across the foundation graphs, specialized queries reveal information hidden by a generic readout, and rigid blade isolation removes cross-grade capacity. We therefore introduce ICE (Interaction-aware Clifford Encoder), a multimodal graph foundation model built on a node-indexed Clifford latent field. Topology, text, and images enter explicit Cl(3) addresses. Edge-aware geometric products transform these directions into scalar, bivector, and trivector relations over observed neighborhoods. A protected Grade-1 route preserves entity semantics, while the full grade and depth bank remains available to fresh node and link heads. We establish exact cross-grade reachability, node-permutation equivariance, and a bound on the task residual around the semantic score. Experiments span one shared foundation over eleven graphs, six node-classification datasets, three link-prediction datasets, and matched few-shot tasks. ICE ranks first in all 30 reported supervised and few-shot comparisons. Core removals reduce every task summary, and mechanism controls connect the gains to higher-order transport, retained multidepth structure, semantic protection, and direct field access.