arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PaSta:基于部分标签学习的噪声节点分类

PaSta: Noisy Node Classification with Partial Label Learning

Yujing Liu, Yixin Liu, Yu Zheng, Yue Tan, Alan Wee-Chung Liew, Shirui Pan

arXiv 2608.25365首次发表:更新:

发表机构

School of Information and Communication Technology, Griffith University(格里菲斯大学信息与通信技术学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对噪声节点分类任务,提出基于部分标签学习的自训练框架PaSta,通过多标注器聚合构建部分标签并结合双损失函数,经闭环迭代优化,在五个数据集上较SOTA方法平均提升1.1%性能。

AI 中文摘要

噪声节点分类是现实世界中与图相关的网络服务的一项基础但极具挑战性的任务,由于弱监督或自动标注,节点标签常常被损坏或不可靠。然而,现有方法通常基于独热标签训练模型,这不仅使模型容易对噪声标签过拟合,还会在伪标签引导的增强后导致误差累积。在本文中,我们提出了一种新颖的基于部分标签的自训练框架(简称 PaSta),它利用部分标签学习技术克服现有方法的局限性。具体而言,PaSta 首先训练多个标注器以全面捕捉节点的类别分布,并聚合它们的预测结果来构建高质量的部分标签。随后,我们设计了一个基于部分标签的分类模型,带有两个精心设计的损失函数,以在标签空间和表示空间两个层面指导模型学习。为进一步增强对噪声标签的鲁棒性,我们引入了一种自训练策略,即通过部分标签学习优化后的标签会以闭环迭代的方式进一步优化标注器。在五个数据集上进行的大量实验表明,与现有最先进的方法相比,PaSta 在各种噪声设置下的分类性能平均提升了 1.1%。

英文摘要

Noisy node classification problem is a fundamental yet challenging task for real-world graph-related web services, where node labels are often corrupted or unreliable due to weak supervision or automatic annotation. However, existing methods typically train models based on one-hot labels, which not only makes models susceptible to overfitting on noisy labels, but also leads to error accumulation after pseudo-label-guided enhancement. In this paper, we propose a novel Partial label-based Self-training framework (PaSta for short) that leverages partial label learning technique to overcome the limitations of existing methods. Specifically, PaSta first trains multiple annotators to comprehensively capture the class distribution of nodes and aggregates their predictions to construct high-quality partial labels. Subsequently, we design a partial label-based classification model with two well-crafted loss functions to guide the model learning at both label and representation spaces. To further enhance the robustness against noisy labels, we introduce a self-training strategy where the labels refined by partial label learning are then used to further optimize the annotators in a closed-loop iterative manner. Extensive experiments on five datasets demonstrate that, compared with existing state-of-the-art methods, PaSta achieves an average improvement of 1.1% in classification performance under various noise settings.

Comments10 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑