arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

时序异构图预训练用于关系深度学习

Temporal Heterogeneous Graph Pretraining for Relational Deep Learning

Yixin Peng, Er Jin, Diego Collarana, Stefan Decker

arXiv 2609.35219首次发表:更新:

发表机构

RWTH Aachen University; Fraunhofer FIT(亚琛工业大学; 弗劳恩霍夫应用信息技术研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对关系深度学习中的时序异构图,提出结合多尺度与旋转时间编码及三目标自监督的两阶段预训练框架,在RelBench数据集上显著提升下游任务性能。

AI 中文摘要

关系深度学习将数据库行和外键链接建模为异构图,以便从记录属性和关系上下文中进行预测。这些图包含两种不同的时间信号:记录年龄随预测截止时间变化,而观察到的记录之间的间隔保持固定。先前的工作通常将时间视为单一信号,或分别研究时序表示和预训练。我们研究了显式编码这两种信号如何影响下游任务的时序预训练。我们的框架结合了多尺度时间编码(通过可学习的时间尺度和类型特定的投影捕获记录年龄)和旋转时间编码(在图传播过程中通过旋转变换表示带符号的记录间间隔)。我们将这些编码与三个自监督目标配对:历史关系恢复、基于时域感知的未来关系活动预测,以及时序子图对比。所有输入都尊重其观察截止时间。预训练分两个阶段进行:子图对比首先学习邻域表示,然后通过关系恢复或未来活动预测进行细化。我们在五个RelBench数据集上,使用异构GNN和图Transformer骨干网络,评估了11个分类和回归任务。在两种编码下,评估的最佳分阶段调度在两种骨干网络上分别比使用相同编码的监督训练提高了3.02%和1.06%,比没有预训练或任一编码的对照组分别提高了3.24%和2.37%。

英文摘要

Relational deep learning models database rows and foreign-key links as a heterogeneous graph for prediction from record attributes and relational context. These graphs contain two distinct temporal signals: record age changes with the prediction cutoff, while intervals between observed records remain fixed. Prior work often treats time as a single signal or studies temporal representation and pretraining separately. We investigate how explicitly encoding both signals affects temporal pretraining for downstream tasks. Our framework combines Multi-scale Time Encoding, which captures record age using learnable time scales and type-specific projections, with Rotary Time Encoding, which represents signed inter-record intervals through rotary transformations during graph propagation. We pair these encodings with three self-supervised objectives: historical relation recovery, horizon-aware future relation activity prediction, and temporal subgraph contrast. All inputs respect their observation cutoffs. Pretraining proceeds in two stages: subgraph contrast first learns neighborhood representations, followed by refinement through either relation recovery or future activity prediction. We evaluate on five RelBench datasets across 11 classification and regression tasks using heterogeneous GNN and graph Transformer backbones. With both encodings, the best evaluated staged schedules improve over supervised training with the same encodings by 3.02% and 1.06% on the two backbones, respectively, and over controls without pretraining or either encoding by 3.24% and 2.37%.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑