发表机构
Australian National University; Data61, CSIRO; University of New South Wales(澳大利亚国立大学; Data61,澳大利亚联邦科学与工业研究组织; 新南威尔士大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对GNN到MLP蒸馏中忽略图几何导致谱欠拟合与过拟合的问题,提出基于Ollivier-Ricci曲率引导的G^2MLP框架,在训练时对齐预测与表示,提升无图推理性能。
AI 中文摘要
GNN到MLP的蒸馏旨在保留消息传递教师模型的预测准确性,同时在推理时部署无图MLP。现有方法主要传递节点级预测或使用基于置信度的重新加权,但未指明学生应在何处保留教师由图诱导的几何结构。我们表明,这一遗漏导致学生表示空间中出现两种谱失败模式。在稀疏图上,学生遭受谱欠拟合,缺失集中在边界区域附近的高能量教师方向。在密集图上,学生遭受谱过拟合,保留教师通过聚合已坍缩的虚假方向。受能量加权教师-学生对齐目标的启发,我们提出图几何感知MLP(G^2MLP),一种由Ollivier-Ricci曲率引导的训练时蒸馏框架。曲率识别两种谱误差集中的位置,并用于在预测级和表示级对齐之间分配监督。部署的模型保持为标准MLP,推理时无需访问图。在节点分类基准上,G^2MLP持续优于无图蒸馏基线,减少两种场景下的教师-学生秩差距,并无需架构更改即可迁移到图Transformer教师和链接预测。
英文摘要
GNN-to-MLP distillation aims to retain the predictive accuracy of a message-passing teacher while deploying a graph-free MLP at inference. Existing methods mainly transfer node-wise predictions or use confidence-based reweighting, but they do not specify where the student should preserve the teacher's graph-induced geometry. We show that this omission leads to two spectral failure modes in the student's representation space. On sparse graphs, the student suffers from spectral underfit, missing high-energy teacher directions concentrated near boundary regions. On dense graphs, it suffers from spectral overfit, retaining spurious directions that the teacher has collapsed through aggregation. Motivated by an energy-weighted teacher-student alignment objective, we propose Graph Geometry-aware MLP (G^2MLP), a training-time distillation framework guided by Ollivier-Ricci curvature. Curvature identifies where the two spectral errors concentrate and is used to allocate supervision between prediction-level and representation-level alignment. The deployed model remains a standard MLP and requires no graph access at inference. Across node-classification benchmarks, G^2MLP consistently improves over graph-free distillation baselines, reduces the teacher-student rank gap in both regimes, and transfers without architectural changes to Graph Transformer teachers and link prediction.