发表机构
Johns Hopkins University; Zhejiang University; National University of Singapore; Grabtaxi Holdings Pte Ltd(约翰斯·霍普金斯大学; 浙江大学; 新加坡国立大学; Grab出租车控股私人有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对图节点分类的类别不平衡问题,提出利用平衡元集评估节点重要性的方法,推导公式并开发新框架,能过滤有价值节点,构建高质量元集,实验证明该框架在缓解不平衡上具优越性。
AI 中文摘要
在实际应用中,图上的节点分类常面临类别不平衡挑战,多数类主导训练导致模型性能有偏差。传统GNN在此场景中表现不佳。现有解决方案存在不足。本文提出利用平衡元集进行重要性度量的方法,识别可抵消类别不平衡的重要节点用于模型训练,理论推导直接评估节点重要性的公式,开发新框架过滤有价值节点,还介绍构建高质量元集的策略。通过实验证明该框架在缓解类别不平衡方面的优越性。
英文摘要
In real-world applications, node classification on graphs often faces the challenge of class imbalance, where majority classes dominate training, resulting in biased model performance. Traditional GNNs often struggle in such scenarios, as they tend to overfit to majority classes while underrepresenting minority classes. Existing solutions, which either prioritize nodes based on class size or synthesize new nodes for minority classes, often fall short of effectively addressing this imbalance issue. This paper introduces an approach to class-imbalanced node classification by utilizing a balanced meta-set for importance measurement, where a training node is considered significant if it enhances model performance under an unbiased setting. Our method identifies important nodes that can counteract class imbalance and utilizes them for model training, allowing for fine-grained and dynamic node selection throughout the training process. We theoretically derive a formula to directly assess node importance, reducing computational overhead and providing an intuitive threshold for node selection. Guided by this metric, we develop a novel framework that filters valuable labeled, unlabeled, and synthetic nodes that enhance model performance in an unbiased context. A key advantage of this framework is its separation of the synthetic node generation process from the filtering process, ensuring compatibility with various node generation methods. Furthermore, we introduce a strategy to construct a high-quality meta-set that closely approximates the overall feature distribution, ensuring robust representation of each class. We evaluate our framework, NodeImport, across multiple datasets using popular GNN architectures, demonstrating its superiority over existing baselines. Our results highlight the flexibility and effectiveness of the framework in mitigating class imbalance, leading to improved outcomes.
Journal refKDD '25: Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1, 2025, Pages 94 - 105