解耦同质性与稀有性:解释图神经网络的失效
Disentangling Homophily and Rarity: Explaining Failure in Graph Neural Networks
浏览论文内容
中文总结 AI 辅助
该研究通过评估6种GNN在5个不同同质性数据集上的表现,挑战了异质节点分类的子群泛化框架,发现可通过重训分类头或最终线性层恢复异质节点分类所需信息。
中文摘要 AI 辅助
图中的异质节点更难分类,是因为它们具有异质性,还是因为它们较为稀有?部分现有研究将这类节点的分类归为子群泛化问题,即模型以牺牲稀有群体的表现为代价,在多数群体上表现良好;另一些研究则将其解释为图神经网络(GNN)的邻域聚合问题。我们通过在5个具有不同同质性程度的数据集上对6种GNN进行详细评估,发现即使同质节点较为稀有,其分类也往往更容易——这对上述子群框架提出了挑战。不过,我们的发现也细化了现有关于GNN如何误表示异质节点的认知:我们证明,正确分类异质节点所需的信息,通常可以通过重新训练模型的分类头,甚至仅重新训练最终的线性分类层来恢复。
英文摘要
Are heterophilic nodes in a graph harder to classify because they are heterophilic or because they are rare? Some existing work frames classification of such nodes as a subgroup generalisation problem, where a model performs well on the majority group at the expense of the rare group. Others explain this as a problem of neighbourhood aggregation in graph neural networks (GNNs). We assess these two viewpoints through a detailed evaluation of six GNNs on five datasets of varying homophily, and find that homophilic nodes tend to be easier to classify, even when they are rare---challenging the subgroup framing. However, our findings also nuance existing beliefs about how GNNs misrepresent heterophilic nodes. We demonstrate that the information needed to classify heterophilic nodes correctly is often recoverable by retraining the classification head of a model, or even just the final linear classification layer.
发表机构
- Simula Research Laboratory(西穆拉研究实验室)
机构由 AI 辅助整理,请以论文原文为准。