arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习潜在结构:一种以特征为中心的图数据增强方法

Learning the Latent Structure: A Feature-Centric Approach to Graph Data Augmentation

Yu Song, Zhigang Hua, Yan Xie, Bingheng Li, Jingzhe Liu, Bo Long, Jiliang Tang, Hui Liu

arXiv 2610.02517首次发表:更新:

发表机构

Michigan State University; Meta(密歇根州立大学; Meta)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出以特征为中心的图数据增强框架SelfAug,通过嵌入空间自监督逆掩码学习潜在结构,在十个数据集上优于现有方法,兼顾准确性与效率。

AI 中文摘要

图结构数据在建模复杂关系中起着关键作用。然而,由于数据收集和观测限制,现实世界的图往往不完整,严重限制了现代图学习流程的有效性。现有的图数据增强(GDA)方法试图优化图结构以提升下游性能,但它们通常依赖标签、计算成本高且本质上是转导式的,限制了其在实际场景中的应用。在本工作中,我们提出了一种新颖的以特征为中心的图数据增强框架,通过在嵌入空间中直接操作来绕过显式结构建模。通过自监督逆掩码过程,我们的方法捕捉观测图与完整图之间的潜在联系,从而通过精炼的节点表示恢复未观测到的结构信号。为了增强在噪声和稀疏监督下的鲁棒性,我们引入了消息正则化器和自举策略,以实现有效训练和泛化。在涵盖多个领域的十个图数据集上评估,我们的方法SelfAug在归纳和冷启动设置中,在准确性和效率方面均持续优于最先进的方法,突显了其作为可扩展且可泛化解决方案在现实世界图学习场景中的潜力。

英文摘要

Graph-structured data plays a pivotal role in modeling complex relationships. However, real-world graphs are often incomplete due to data collection and observational constraints, severely limiting the effectiveness of modern graph learning pipelines. While existing Graph Data Augmentation (GDA) methods attempt to refine graph structures for improved downstream performance, they are typically label-dependent, computationally expensive, and inherently transductive, limiting their applicability in practical scenarios. In this work, we present a novel feature-centric graph data augmentation framework that bypasses explicit structure modeling by operating directly in the embedding space. Through a self-supervised inverse masking process, our method captures latent ties between observed and complete graphs, enabling recovery of unobserved structural signals through refined node representations. To enhance robustness under noisy and sparse supervision, we introduce a message regularizer and a bootstrap strategy for effective training and generalization. Evaluated on ten graph datasets spanning multiple domains, our approach, SelfAug, consistently outperforms state-of-the-art methods in both accuracy and efficiency across inductive and cold-start settings, highlighting its potential as a scalable and generalizable solution for real-world graph learning scenarios.

CommentsAccepted at AAAI 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑