发表机构
Universität Bayreuth; FAU Erlangen-Nürnberg; Ludwig-Maximilians-Universität München; Munich Center for Machine Learning(拜罗伊特大学; 埃尔朗根-纽伦堡大学; 慕尼黑大学; 慕尼黑机器学习中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对图神经网络评估中随机划分导致结果不稳定的问题,提出同质性感知分层程序,以节点同质性为主轴、类别为次轴构建划分,在15个数据集和7种架构上取得更低的稳定性排名。
AI 中文摘要
图神经网络被广泛用于转导式节点分类,其准确率通常是在随机划分的训练/验证/测试集上测量的。已有研究表明,对于同一数据集的不同随机划分,报告的准确率会显著变化,这使得已发表的架构间比较变得不可靠。在非图场景中,经典的解决方法是分层k折交叉验证,它确保每个测试折都反映数据集的完整类别分布。我们认为,仅进行类别分层对于图来说是不够的:节点不是孤立的,而是相互连接的,不同折在局部邻域同质性分布上的差异会使模型暴露于系统性不同的关系条件下,这些条件直接影响消息传递行为。由此产生的跨折变异反映了每个划分的同质性组成,使得报告方差高于仅由模型行为本身所产生的方差。为解决这一问题,我们提出了\hp{},一种拓扑感知的分层程序,它将节点同质性作为主要分层轴,在标准分层已控制的类别边际之外,使各折在局部关系一致性上对齐。仅按同质性分层并不能保证类别平衡,因此\hp{}将类别标签作为次要轴,从而自然保持了类别的代表性。我们在一个广泛的基准套件上评估\hp{},该套件包含15个跨越完整同质性谱系的节点分类数据集和7种GNN架构。\hp{}实现了1.49的平均稳定性排名,而随机k折为2.31,在15个数据集中的13个上取得了最低的平均稳定性排名,同时保持了接近类别分层划分的类别平衡,且显著优于随机划分。我们认为,同质性感知的划分构建值得在GNN评估中更广泛地采用。
英文摘要
Graph neural networks are widely used for transductive node classification, with accuracy typically measured on randomly drawn train/validation/test splits. Reported accuracy has been shown to shift substantially across different random splits of the same dataset, making published comparisons between architectures unreliable. The classical remedy in non-graph settings is stratified $k$-fold cross-validation, which ensures each test fold reflects the full class distribution of the dataset. We argue that class stratification alone is insufficient for graphs: nodes are not isolated but connected, and folds that differ in their distribution of local neighbourhood homophily expose the model to systematically different relational conditions that directly affect message-passing behaviour. The resulting cross-fold variation reflects the homophily composition of each split, inflating reported variance beyond what model behaviour alone would produce. To address this, we propose \hp{}, a topology-aware stratification procedure that treats node homophily as the primary stratification axis, aligning folds with respect to local relational consistency alongside the class marginal that standard stratification already controls. Stratifying on homophily alone does not guarantee class balance, so \hp{} incorporates class label as a secondary axis, preserving class representativeness as a natural consequence of the procedure. We evaluate \hp{} on a broad benchmark suite comprising 15 node-classification datasets spanning the full homophily spectrum and 7 GNN architectures. \hp{} achieves a mean stability rank of 1.49 compared to 2.31 for random $k$-fold, achieving the lowest mean stability rank on 13 of 15 datasets while preserving class balance close to class-stratified splits and substantially better than random. We argue that homophily-aware split construction merits broader adoption for GNN evaluation.
Comments10 pages