arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OpenRTAG:数据质量退化下鲁棒文本属性图学习的综合基准

OpenRTAG: A Comprehensive Benchmark for Robust Text-Attributed Graph Learning under Data Quality Degradation

Yuze Dai, Zhihan Zhang, Yan Zhao, Ruoyu Wu, Xunkai Li, Zekai Chen, Qiangqiang Dai, Hongchao Qin, Ronghua Li

arXiv 2607.19108首次发表:更新:

AI 中文总结

针对现实中TAGs存在质量问题影响学习的情况,提出OpenRTAG基准,将质量问题分类,支持多数据集和任务的标准化评估,能系统评估场景与模型,比较多种模型,研究基线情况及复合场景下模型行为,为理解TAG学习鲁棒性提供平台。

AI 中文摘要

文本属性图(TAGs)是一种重要的图数据形式,结合了关系结构与丰富的节点文本。然而,现实世界中的TAGs往往不完美,存在文本、结构和标签方面的质量问题,通常表现为稀疏性、噪声和不平衡。这些维度定义了九种代表性的退化场景,会严重影响TAG学习。虽然先前研究探索了特定缓解策略,但现有证据在退化类型、数据集、任务和模型家族中仍很零散。为填补这一空白,我们提出OpenRTAG,一个文本属性图学习的鲁棒性基准。它将TAG质量问题组织成统一的3*3分类法,支持在九个TAG数据集和三个下游任务上进行标准化评估。系统评估场景有效性和模型敏感性,比较传统GNN、LLM - GNN和代表性GFM,研究场景匹配基线的有效性、效率和鲁棒性,并进一步检查复合退化场景下的模型行为。OpenRTAG提供了一个标准化测试平台,用于理解现实低质量设置下TAG学习中的鲁棒性。

英文摘要

Text-attributed graphs (TAGs) are an important graph data form that combine relational structure with rich node text. However, real-world TAGs are often imperfect, with quality issues arising from text, structure, and labels, and typically manifesting as sparsity, noise, and imbalance. These dimensions define nine representative degradation scenarios that can substantially affect TAG learning. Although prior studies have explored specific mitigation strategies, existing evidence remains fragmented across degradation types, datasets, tasks, and model families, leaving TAG robustness insufficiently understood. To address this gap, we present OpenRTAG, a robustness benchmark for text-attributed graph learning. OpenRTAG organizes TAG quality issues into a unified 3 * 3 taxonomy and supports standardized evaluation across nine TAG datasets and three downstream tasks. It systematically evaluates scenario validity and model sensitivity, compares traditional GNNs, LLM-GNNs, and a representative GFM, investigates the effectiveness, efficiency, and robustness of scenario-matched baselines, and further examines model behavior under composite degradation scenarios. OpenRTAG provides a standardized testbed for understanding robustness in TAG learning under realistic low-quality settings.

Comments9 pages, 10 figures, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑