arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于图神经网络的归纳式关联聚类

Inductive Correlation Clustering with Graph Neural Networks

Francesco Paolo Nerini, Francesco Bonchi, Arijit Khan, André Panisson

arXiv 2608.27153首次发表:更新:

发表机构

Sapienza University of Rome; Intesa Sanpaolo AI Research; Bowling Green State University(罗马大学; 裕信银行人工智能研究院; 鲍灵格林州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对关联聚类算法的直推式缺陷,提出基于图神经网络的归纳式关联聚类框架,可高效推广至新图,在归纳设置下推理时间大幅降低,还可作为图分类的高效池化层。

AI 中文摘要

关联聚类(Correlation Clustering,CC)是组合优化中一种自然的聚类形式,它采用输入的图表示,无需预先指定聚类数量。给定n个对象和一个成对相似度函数,目标是将对象聚类,使相似对象置于同一聚类,不相似对象置于不同聚类。尽管CC具有通用性,但现有CC算法存在严重的可扩展性问题,且本质上是直推式的:即算法必须针对任何新问题实例从头执行。在本研究中,我们利用图神经网络(Graph Neural Networks,GNN)弥合这一差距,以解决归纳式关联聚类问题——这是为处理未见图实例而设计的CC问题的新推广。通过在训练过程中学习利用常见结构模式和节点特征,我们的框架可推广到来自同一分布的新图,且与标准算法相比计算开销极小。我们通过大量实验证明了所提方法的有效性和可扩展性:我们的框架不仅在归纳式设置中表现出色,例如将推理时间降低多达5个数量级,同时保持近似率在最佳基准解的10%以内,还在标准(直推式)CC基准上取得了有竞争力的结果。最后,我们展示了该框架作为图分类可池化机制的实际应用,结果表明我们的方法可作为高效池化层,增强GNN捕捉网络中分层结构信息的能力。

英文摘要

Correlation Clustering (CC) is a natural formulation of clustering in combinatorial optimization, which uses a graph representation of the input and does not require a pre-specified number of clusters. Given $n$ objects and a pairwise similarity function, the goal is to cluster the objects so that similar objects are put in the same cluster and dissimilar objects are put in different clusters. Despite its versatility, existing CC algorithms suffer from significant scalability issues and are inherently transductive: i.e., the algorithm must be executed from scratch for any new problem instance. In this work, we bridge this gap by leveraging Graph Neural Networks (GNNs) to solve Inductive Correlation Clustering, a novel generalization of the CC problem designed to handle unseen graph instances. By learning to exploit common structural patterns and node features during training, our framework generalizes to new graphs drawn from the same distribution with minimal computational overhead with respect to standard algorithms. We demonstrate the effectiveness and scalability of our approach through extensive experiments. Our framework not only excels in the inductive setting, e.g., lowering the inference time up to $5$ order of magnitude, while maintaining an approximation ratio within $~10\%$ of the best baseline solution, but also achieves competitive results on standard (transductive) CC benchmarks. Finally, we showcase a practical application of our framework as a learnable pooling mechanism for graph classification. Our results indicate that our method serves as an efficient pooling layer, enhancing the ability of GNNs to capture hierarchical structural information in networks.

CommentsAccepted at CIKM'26

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑