AI 中文总结
针对异配图节点分类,提出亲和力引导的图重连方法,联合估计图与嵌入,在多个基准上提升准确率并降低分类器差异。
AI 中文摘要
图神经网络在异配图上会失去很大一部分优势,在异配图中,相连的节点往往携带不同的标签。图重连是一种流行的补救措施,但重连方法通常使用单一分类器进行评估,这使得很难判断增益是来自新的拓扑结构还是来自特定的分类器配对。我们提出了一种亲和力引导的重连方法,该方法同时估计图和节点表示。它遵循期望最大化的精神,在基于当前图训练一个轻量级图神经网络与在模块度目标下结合伪标签同质性项对候选边进行重新加权之间交替进行。候选边来自一个紧凑的候选池,该候选池通过对比学习得到的节点相似度和邻域分布亲和力进行评分。该方法返回两个与分类器无关的输出:一个重连后的图和一个在该图上学习到的节点嵌入。在六个异配基准和五个下游分类器上,与使用归一化特征的原始图相比,该方法在30个分类器-数据集组合中的23个上提高了准确率,平均增益为5.8个百分点,并将分类器之间的准确率离散度降低了约四倍。一项受控消融研究表明,这两个输出各自有用且发挥互补作用:嵌入贡献了大部分准确率增益,而重连后的图使不同分类器趋于一致。一个完全无监督的变体在重连过程中不使用标签,保留了大部分改进。重连后的图也更具同质性,并改善了标签传播和社区检测。
英文摘要
Graph neural networks lose much of their advantage on heterophilic graphs, where connected nodes often carry different labels. Graph rewiring is a popular remedy, but rewiring methods are usually evaluated with a single classifier, which makes it hard to tell whether the gains come from the new topology or from that particular pairing. We propose an affinity-guided rewiring method that estimates the graph and the node representation together. It alternates, in the spirit of expectation maximisation, between training a lightweight graph neural network on the current graph and re-weighting candidate edges under a modularity objective with a pseudo-label homophily term. Candidate edges come from a compact pool scored by a contrastively learned node similarity and a neighbourhood-distribution affinity. The method returns two classifier-independent outputs: a rewired graph and a node embedding learned on it. Across six heterophilic benchmarks and five downstream classifiers, it improves accuracy over the original graph with normalised features in 23 of 30 classifier-dataset combinations, with a mean gain of 5.8 points, and reduces the accuracy spread between classifiers about fourfold. A controlled ablation shows that the two outputs are each useful and play complementary roles: the embedding contributes most of the accuracy gain, while the rewired graph makes different classifiers agree. A fully unsupervised variant, which uses no labels during rewiring, retains most of the improvement. The rewired graphs are also more homophilic and improve label propagation and community detection.