arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

时空预测基准数据集与基线模型的批判性审查

A Critical Audit of Spatiotemporal Forecasting Benchmark Datasets and Models

Kenneth Martin, Simon Heilig, Asja Fischer, Michel F. C. Haddad, Adam M. Sykulski, Moshe Eliasof

arXiv 2608.20980首次发表:更新:

发表机构

Imperial College London; Ruhr University Bochum; Queen Mary University of London; Ben-Gurion University of the Negev(帝国理工学院; 鲁尔大学波鸿分校; 伦敦玛丽女王大学; 内盖夫本-古里安大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文审查了时空预测领域的基准数据集与基线,发现空间无关线性模型表现优于预期,揭示了一阶差分数据集的结构偏差,建议减少对这类数据集的依赖并提出了应用分析结果开发GNN模型的新途径。

AI 中文摘要

图神经网络(GNN)常被用于具有空间图结构的多元时间序列的短期预测。尽管存在许多替代数据集,该领域的方法创新主要针对一组相当有限的基准数据集进行评估,最著名的是水痘(Chickenpox)、PedalMe、WikiMaths、METR-LA和PEMS-BAY。评估协议包含从历史平均到经典机器学习方法的基线,这些基线通常表现出与GNN相比具有竞争力的性能。在本研究中,我们退一步通过经典时间序列方法分析基准数据集,以揭示为什么空间无关的线性模型比之前报道的更具竞争力,这进一步对上述广泛采用的数据集的判别可靠性提出了质疑。我们的统计分析提供了一套工具,用于识别显著的空间和时间相关性,同时揭示了一阶差分数据集引入的结构偏差。因此,我们建议减少对这类数据集在方法比较中的过度依赖,转而倡导更严格的统计评估。通过将我们的分析结果应用于一个简单的混合模型,我们展示了我们的方法如何能为开发GNN模型带来新的途径。

英文摘要

Graph neural networks (GNNs) are routinely employed for spatiotemporal forecasting, yet their performance across widely used benchmark datasets is inconsistent. Here, we perform an audit of dataset properties and baseline models to assess the quality of the benchmarks, and the robustness of the conclusions drawn from them. Using classical statistical tools, we characterise spatiotemporal lagged dependencies in benchmarks, and examine how temporal differencing changes these relationships and affects model rankings. Motivated by this, we re-evaluate temporal linear baselines, significantly reducing the apparent gains from GNNs on several benchmarks, and surpassing GNNs on others. Suspecting that GNNs struggle to extract linear, node-wise signals, we find that supplying them with autoregressive residuals improves their performance particularly on non-traffic benchmarks. Finally, controlled synthetic experiments reveal that GNNs are sensitive to heterogeneity in temporal dynamics and spatial graph interactions. Together, our findings demonstrate that baseline specification, data pre-processing and system heterogeneity shape the interpretations drawn from benchmark rankings, informing the design and robust evaluation of GNNs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑