arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

半监督文本属性图蒸馏

Semi-Supervised Text-Attributed Graph Distillation

Yurui Lai, Samir Moustafa, Renchi Yang, Tsz Nam Chan

arXiv 2607.20477首次发表:更新:

发表机构

Hong Kong Baptist University; CeMM Research Center for Molecular Medicine of the Austrian Academy of Sciences; Universität Vienna; Shenzhen University(香港浸会大学; 奥地利科学院分子医学CeMM研究中心; 维也纳大学; 深圳大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对文本属性图表示学习的可扩展性瓶颈等问题,提出基于Wasserstein距离的半监督框架\algo{},通过图-文本协作编码模块等方法,在基准数据集实验中实现了高效的TAG学习及性能-压缩权衡。

AI 中文摘要

文本属性图(TAGs)已成为一种用于整合图拓扑与丰富文本语义的表达性数据模型。现有TAGs表示学习方法存在严重的可扩展性瓶颈,尤其是与大语言模型(LLMs)结合时。数据蒸馏提供了一种有前景的以数据为中心的解决方案,但现有方法未能捕捉图与文本模态间的复杂相互作用,难以应对半监督设置中固有的标签稀缺问题,且缺乏生成下游基于LLM任务所需的人类可读文本属性的能力。为应对这些挑战,我们提出了基于Wasserstein距离(WSD)的统一半监督框架\algo{}。它引入了图-文本协作编码模块,利用协作自训练方案中的双路径编码器(图感知和无图)来获取可靠的伪标签并融合互补的图-文本特征。此外,还开发了基于WSD的理论基础图草图算法和具有成本效益的LLM文本合成模块。在基准数据集上的广泛实验表明,\algo{}在基于GNN和LLM的下游任务中实现了最先进的性能-压缩权衡,实现了高效的TAG学习或分析。

英文摘要

{\em Text-Attributed Graphs} (TAGs) have emerged as an expressive data model for integrating graph topology with rich textual semantics. Existing representation learning methods over TAGs suffer from severe scalability bottlenecks, particularly together with {\em Large Language Models} (LLMs). While data distillation offers a promising data-centric solution, existing methods fail to capture the complex interplay between graph and text modalities, struggle with the label scarcity inherent in semi-supervised settings, and lack the ability to produce the human-readable textual attributes required for downstream LLM-based tasks. To address these challenges, we propose \algo{}, a unified semi-supervised framework guided by the {\em Wasserstein Distance} (WSD). Grounded in our empirical findings on real TAGs, \algo{} introduces a graph-text collaborative encoding module that utilizes dual-pathway encoders (graph-aware and -free) within a collaborative self-training scheme to harvest reliable pseudo-labels and fuse complementary graph-text features. Furthermore, we develop a theoretically grounded WSD-based graph sketching algorithm and a cost-effective LLM text synthesis module, which leverages cluster-based keyword extraction to generate coherent, human-readable summaries for condensed nodes. Extensive experiments on benchmark datasets demonstrate that \algo{} achieves a state-of-the-art performance-compression trade-off in terms of both GNN- and LLM-based downstream tasks, enabling effective and efficient TAG learning or analytics.

CommentsTechnical report for the paper "Semi-Supervised Text-Attributed Graph Distillation" accepted KDD2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑