发表机构
University of Westminster(威斯敏斯特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究复现道路级事故预测的图神经网络,发现其多数设计决策无法复现,且与按历史事故数排序的无参数基线相比无显著优势,表明报告的性能提升多为短视野基线造成的假象。
AI 中文摘要
图神经网络越来越多地被应用于道路级事故预测,但其报告的性能提升的稳定性却鲜少受到审视。我们独立重建了一个近期提出的不确定性感知模型的数据管道,并在三个伦敦行政区、采用扩展窗口协议下评估了其十一项设计决策。其中四项在第二个行政区上复现成功,七项未能复现,且其中四项的效应方向发生反转而非仅幅度减弱。多随机种子评估具有决定性:一个效应在单个行政区内不同随机种子间发生符号反转,而参考架构的每行政区种子差异高达35.7个百分点,相比之下我们的模型仅为4个百分点。我们进一步将两个网络与一个无参数基线进行比较,该基线按历史累积事故数对路段进行排序。在匹配的历史深度下,我们的模型与该基线在统计上无显著差异(-0.90个百分点,p=0.61),而参考架构在18个保留窗口上全部输给该基线(-17.37,p<10^{-6})。对基线的回溯视野进行扫描显示,仅该变量即可解释22.71%至83.94%的准确率变化,并且该研究领域中的每一个已发表数值都能被基线在1至5年的视野下匹配。我们认为,图网络在该任务上相对于历史基线的表面优势,很大程度上是基线计算视野过短造成的假象,并建议将视野匹配的基线和多随机种子报告作为最低实践要求。
英文摘要
Graph neural networks are increasingly applied to road-level crash prediction, but the stability of their reported gains has received little scrutiny. We independently reconstruct the data pipeline of a recent uncertainty-aware model and evaluate eleven of its design decisions across three London boroughs under an expanding-window protocol. Four survive replication on a second borough; seven do not, and four of those reverse sign rather than attenuate. Multi-seed evaluation is decisive: one effect reverses sign between random seeds within a single borough, and the reference architecture exhibits per-borough seed spreads of up to 35.7 points against 4 points for ours. We further compare both networks against a parameter-free baseline that ranks segments by cumulative past crash count. At matched history depth our model is statistically indistinguishable from that baseline ($-0.90$ points, $p=0.61$), and the reference architecture loses to it on 18 of 18 held-out windows ($-17.37$, $p<10^{-6}$). Sweeping the baseline's lookback horizon shows it spans 22.71% to 83.94% accuracy on that variable alone, and that every published figure in this line of work is matched by the baseline at a horizon of one to five years. We argue that the apparent margin of graph networks over historical baselines in this task is substantially an artefact of the short horizons those baselines were computed over, and recommend horizon-matched baselines and multi-seed reporting as minimum practice.
Comments6 pages, 1 figure, 4 tables. Code and reproduction scripts: https://github.com/Maurya1112-sudo/greyspot