发表机构
CNRS; IRISA; INRIA(法国国家科学研究中心; IRISA研究所; 法国国家信息与自动化研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
CoRe-GNN是一种基于粗化图的多级消息传递图神经网络,并行执行粗化聚类间与局部聚类内传播,在多种节点分类基准上优于图粗化和Cluster-GCN,兼具长程任务竞争力与内存效率。
AI 中文摘要
在大图上训练图神经网络(GNN)面临着各层存储所有节点表示的内存成本挑战。我们表明,几种现有可扩展方法可被写为GNN传播矩阵的结构化修改,提供了统一视角以揭示它们各自的局限性。具体而言,图粗化用低秩近似替换传播矩阵,该近似可提供谱保证,但会为聚类节点分配统一表示;而Cluster-GCN将传播矩阵限制在聚类内连接,允许高效批处理,但会切断长程信息。这些是图分解为节点组的同一分解的互补性失败。为兼顾两者优势,我们提出CoRe-GNN,它在每一层并行执行两种传播:捕获长程结构的粗化聚类间项,以及保留每个节点判别性的局部聚类内项。我们证明CoRe-GNN继承了与图粗化类似的近似保证,并引入了自然的基于聚类的批处理方案,该方案可扩展到具有数百万节点的图。在涵盖同质性、异质性、大规模及长程图的节点分类基准上,CoRe-GNN的性能优于图粗化和Cluster-GCN基线。值得注意的是,CoRe-GNN在长程任务上达到了有竞争力的准确率,同时通过批处理保持内存效率。
英文摘要
Training Graph Neural Networks on large graphs is challenged by the memory cost of storing all node representations across layers. We show that several existing scalable approaches can be written as structured modifications of the GNN propagation matrix, providing a unified perspective that exposes their respective limitations. In particular, graph coarsening replaces it by a low-rank approximation that enables spectral guarantees but assigns uniform representations to clustered nodes, while Cluster-GCN restricts the propagation matrix to intra-cluster connections that allow efficient batching but sever long-range information. These are complementary failures of the \emph{same} decomposition of the graph into groups of nodes. To obtain the best of both worlds, we propose \textbf{CoRe-GNN}, which performs both propagations in parallel at each layer: a coarsened inter-cluster term capturing long-range structure, and a local intra-cluster term preserving per-node discriminability. We prove that CoRe-GNN inherits analogous approximation guarantees to those of graph coarsening, and introduce a natural cluster-based \emph{batching scheme} that scales to graphs with millions of nodes. On node classification benchmarks spanning homophilic, heterophilic, large-scale, and long-range graphs, CoRe-GNN outperforms both graph coarsening and Cluster-GCN baselines. Notably, CoRe-GNN reaches competitive accuracy on \emph{long-range} tasks, while remaining memory-efficient through batching.