发表机构
Czech Technical University in Prague(布拉格捷克技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对关系深度学习(RDL)现有评估忽略数据随时间演化的局限,提出增量多轮评估与训练范式,发现多数RDL任务存在时间概念漂移,增量微调模型性能优于从头训练基线。
AI 中文摘要
关系深度学习(Relational Deep Learning, RDL)将多表数据库建模为时间异质图,以实现端到端表示学习。然而,主流RDL评估实践依赖静态的单轮数据集快照,忽略了现实数据库持续、随时间演化的特性,导致当前RDL基准无法捕捉模型性能随新数据随时间累积的变化。为解决这一局限,本文提出一种增量多轮评估与训练范式,以评估并提升最先进RDL模型的时间鲁棒性与适应性。利用已有的大规模数据集,本文研究了数据演化与模型训练动态,发现多数预测任务中存在时间概念漂移;本文提出多种用于模型微调的增量训练机制,证明迁移学习在RDL场景中既可行又高效。结合一种优先考虑近期未来准确性的新型时间评估指标,本文表明,经增量微调的模型始终优于标准的、成本高昂的从头训练基线模型。
英文摘要
Relational Deep Learning (RDL) models multi-tabular databases as temporal heterogeneous graphs to enable end-to-end representation learning. However, prevailing RDL evaluation practices rely on static, single-episode dataset snapshots, overlooking the continuous, time-evolving nature of real-world databases. Consequently, current RDL benchmarks fail to capture how model performance changes as new data accumulates over time. To address this limitation, we introduce an incremental, multi-episode evaluation and training paradigm to assess and improve the temporal robustness and adaptability of state-of-the-art RDL models. Using established large-scale datasets, we examine data evolution and model training dynamics, demonstrating that temporal concept drifts occur in the majority of predictive tasks. We present multiple incremental training regimes for fine-tuning the models and demonstrate that transfer learning is both feasible and highly effective in the RDL setting. Alongside a new temporal evaluation metric that prioritizes near-future accuracy, we show that our incrementally fine-tuned models consistently outperform the standard, expensive, from-scratch trained baselines.