发表机构
Zhejiang University(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出TTGBench基准,联合评估文本属性时序图中的结构演化与语义漂移,首个支持多类多标签时序节点分类,揭示TGNN与LLM范式的能力鸿沟。
AI 中文摘要
时序图学习对动态系统的演化进行建模,其中结构交互和语义状态均随时间变化。然而,现有基准主要通过时序链接预测(TLP)强调结构演化,而对语义演化的支持仍然有限。尽管有时会包含时序节点分类(TNC),但它通常局限于过于简单的二分类设置,无法捕捉真实的语义漂移。此外,常用数据集表现出较高的链接重复率,导致性能估计虚高,掩盖了模型的真实能力。为解决这些局限性,我们提出了TTGBench,这是一个联合评估结构演化与语义演化的新基准。TTGBench包含六个具有“双重波动性”特征的真实世界文本丰富数据集,能够对现有模型进行严格且公平的评估。值得注意的是,它是首个同时支持多类和多标签TNC的基准,填补了时序语义漂移评估中的关键空白。我们对17种最先进的方法进行了全面评估,涵盖时序图神经网络(TGNN)和基于大语言模型(LLM)的范式。结果揭示了两种范式之间明显的“能力鸿沟”:基于TGNN的方法在结构预测方面表现出色,但在语义追踪方面失败,而基于LLM的预测器则呈现相反趋势。通过深入分析,我们揭示了它们的基本局限性,并为开发更全面的时序图模型提供了见解。
英文摘要
Temporal graph learning models the evolution of dynamic systems, where both structural interactions and semantic states change over time. However, existing benchmarks primarily emphasize structural evolution via temporal link prediction (TLP), while support for semantic evolution remains limited. Although temporal node classification (TNC) is sometimes included, it is typically restricted to simplistic binary settings that fail to capture realistic semantic drift. Moreover, commonly used datasets exhibit high link repetition, leading to inflated performance estimates and obscuring true model capability. To address these limitations, we introduce \textbf{TTGBench}, a new benchmark that jointly evaluates structural and semantic evolution. TTGBench comprises six real-world, text-rich datasets characterized by \emph{Dual Volatility}, enabling rigorous and fair evaluation of existing models. Notably, it is the first benchmark to support both multi-class and multi-label TNC, filling a critical gap in evaluating temporal semantic drift. We conduct a comprehensive evaluation of 17 state-of-the-art methods across Temporal Graph Neural Networks (TGNNs) and Large Language Model (LLM)-based paradigms. The results reveal a clear \emph{capability divide} between the two paradigms: TGNN-based methods excel at structural prediction but fail at semantic tracking, whereas LLM-based predictors show the opposite trend. Through in-depth analysis, we uncover their fundamental limitations and provide insights for developing more comprehensive temporal graph models.
Comments24 pages, 8 figures, 22 tables