AI 中文总结
提出OFAG基础模型,通过合成数据训练单一模型直接应用于多样属性图聚类,无需针对特定图训练或调优,在十个数据集上以12.43分钟完成全部任务,性能优于现有方法。
AI 中文摘要
属性图聚类旨在通过联合利用节点属性和图拓扑来发现节点群体,然而其无监督特性使得模型选择和适应 inherently 困难。现有方法通常针对每个输入图训练和调优单独的模型,导致流程成本高昂且脆弱,往往无法在不同特征空间、结构模式和属性-结构相关性的图之间迁移。本文研究一种“一劳永逸”的替代方案:能否训练一个单一模型,直接应用于多样化的属性图,而无需针对特定图进行训练、微调或超参数搜索?我们提出 OFAG,一种用于属性图聚类的基础模型。基于先验数据拟合网络(Prior-data Fitted Networks),OFAG 从在潜在簇、节点属性和图结构的广泛先验下生成的合成属性图中学习可复用的聚类推理策略。为处理图间不兼容的特征空间,OFAG 采用维度无关的信号级图编码器,将每个特征通道视为图信号,并建模其对共享图滤波器的响应。模型使用超球面聚类目标进行训练,在推理时通过单次前向传播生成聚类友好的节点表示。在十个数据集上,一个冻结的 OFAG 模型在 NMI、ACC、ARI 和 F1 上取得了最佳平均性能和平均排名,同时完成全部十个数据集仅需总计 12.43 分钟——比第二快的基线快 6 倍以上,比聚类质量第二好的方法快近 28 倍。我们的代码和预训练检查点可在该 https URL 获取,使从业者无需额外训练或调优即可将 OFAG 直接应用于自己的属性图数据集。
英文摘要
Attributed graph clustering aims to discover node groups by jointly exploiting node attributes and graph topology, yet its unsupervised nature makes model selection and adaptation inherently difficult. Existing methods typically train and tune a separate model for each input graph, leading to costly and fragile pipelines that often fail to transfer across graphs with different feature spaces, structural patterns, and attribute-structure correlations. In this paper, we study a one-for-all alternative: can a single model be trained once and directly applied to diverse attributed graphs without graph-specific training, fine-tuning, or hyperparameter search? We propose OFAG, a foundation model for attributed graph clustering. Building upon Prior-data Fitted Networks, OFAG learns a reusable clustering inference strategy from synthetic attributed graphs generated under broad priors over latent clusters, node attributes, and graph structures. To handle incompatible feature spaces across graphs, OFAG adopts a dimension-agnostic signal-wise graph encoder that treats each feature channel as a graph signal and models its response to shared graph filters. The model is trained with a hyperspherical clustering objective, producing clustering-friendly node representations in a single forward pass at inference time. On ten datasets, one frozen OFAG model achieves the best mean performance and average rank across NMI, ACC, ARI, and F1, while completing all ten datasets in 12.43 minutes total---over 6* faster than the second-fastest baseline and nearly 28* faster than the second-best on clustering quality. Our code and pretrained checkpoint are available at https://github.com/Cloudy1225/OFAG, allowing practitioners to directly apply OFAG to their own attributed graph datasets without additional training or tuning.