为何图学习无法充分受益于文本教师?
Why Does Graph Learning Fail to Fully Benefit from a Text Teacher?
浏览论文内容
中文总结 AI 辅助
该研究针对结合自监督GNN与交替优化语言模型的多模态图学习模型未充分提升性能的问题,分析了六个关键影响因素并通过分阶段实验验证。
中文摘要 AI 辅助
图神经网络(GNN)被广泛用于表示实体间的复杂交互与关系。我们研究了一种结合两种互补思路的多模态模型:一是自监督方法,使在某一数据集上预训练的GNN编码器可直接在节点特征维度不同的另一数据集上运行,无需重建模型或重新对齐数据;二是交替优化方法,在E步更新语言模型模块,在M步更新GNN模块,而非端到端联合训练大型语言模型与GNN以处理大图。尽管预期如此,该组合模型并未充分提升预测性能。我们确定了六个因素:(1)E步的外部锚点存在强度-安全权衡:弱锚点几乎无影响,而过强的锚点会损害图表示;(2)E步教师的知识未直接注入GCN嵌入Z;(3)M步构建的表示空间未针对与E步教师空间相同的目标优化,导致为目标分类生成折衷表示;(4)GCN传播会将节点自身的文本信息与邻居信息平均;(5)余弦对齐无法保证对分类有判别力的轴,因此与E步文本锚点更强的几何对齐未必能充分改善目标决策边界或分类性能;(6)M步中保留源侧自监督几何的力,与将表示推向E步教师的力存在冲突。我们通过一系列改变E步影响程度的分阶段实验验证了这些观察结果。
英文摘要
Graph neural networks (GNNs) are widely used to represent complex interactions and relationships among entities. We investigate a multimodal model that combines two complementary ideas: a self-supervised method that enables a GNN encoder pretrained on one dataset to operate directly on another dataset with a different node-feature dimensionality, without rebuilding the model or realigning the data; and an alternating optimization method that updates a language-model module in an E-step and a GNN module in an M-step, rather than jointly training a large language model and a GNN end to end on a large graph. Despite expectations, the combined model did not sufficiently improve predictive performance. We identify six factors: (1) an external anchor in the E-step has a strength-safety trade-off: a weak anchor has little effect, whereas an overly strong anchor can damage the graph representation; (2) the knowledge of the E-step teacher is not injected directly into the GCN embedding Z; (3) the representation space constructed in the M-step is not optimized for the same objective as the E-step teacher space, resulting in a compromise representation for target classification; (4) GCN propagation averages a node's own textual information with information from its neighbors; (5) cosine alignment does not guarantee axes that are discriminative for classification, so stronger geometric alignment with the E-step text anchor need not sufficiently improve the target decision boundary or classification performance; and (6) the force that preserves the source-side self-supervised geometry in the M-step conflicts with the force that moves the representation toward the E-step teacher. We support these observations through a staged set of experiments that varies the influence of the E-step.
发表机构
- SOKENDAI(总研究大学院大学)
- National Institute of Informatics(情报学研究所)
机构由 AI 辅助整理,请以论文原文为准。