arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.12462cs.AI

交通预测中全局空间信息提取真的需要Transformer吗?

Do We Really Need Adaptive Global Spatial Attention for Traffic Forecasting?

Qihang Zhang, Siyao Zhang, Letao Kang, Wenzhe Liang, Miao Zhang, Zhao Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

研究交通预测中全局空间信息提取,设计可控消融框架,对比均匀全范围混合和标准空间注意力,发现二者MAE相近,前者降低空间混合复杂度,机制分析分解空间注意力,揭示残差特性及空间注意力合理性,代码公开。

中文摘要 AI 辅助

现有交通预测模型通常专注于提取空间依赖性,尤其是全局空间信息,其通过交通网络中各节点间交互来表征。但全局信息建模与提取的潜在机制研究不足。不清楚全局信息是否必须通过高自由度自适应注意力提取,还是能用简单全局聚合算子捕获。为此设计可控消融框架,仅替换空间混合模块测试基于注意力的全局交互。在六个交通基准测试中,均匀全范围混合和标准空间注意力在三个数据集上MAE更低,均值MAE仅差0.14%,前者将节点级空间混合复杂度从O(N2)降至O(N)。机制分析将空间注意力分解为行均匀全局背景和非均匀残差。残差显示出数据集相关的边际价值,表明空间注意力应通过超出行均匀全局背景的稳定增益来证明其合理性。相应源代码可公开获取。

英文摘要

Existing traffic forecasting models commonly focus on extracting spatial dependencies, particularly global spatial information, which characterizes the representations obtained through interactions between each node and all nodes across the traffic network. However, the underlying mechanism by which global information is modeled and extracted remains insufficiently investigated. Whether global information must be extracted by high-degree-of-freedom adaptive attention or can be captured by a simple global aggregation operator remains unclear. For this purpose, we design a controlled ablation framework that replaces only the spatial mixing module to test attention-based global interaction. Across six traffic benchmarks, standard spatial attention yields relative MAE changes of $-1.58\%$ to $+1.26\%$ compared with uniform full-range mixing, and we observe no consistent advantage for standard spatial attention, while uniform full-range mixing reduces node-scale spatial-mixing complexity from $O(N^2)$ to $O(N)$. We further propose a hypothesized model that decomposes spatial attention into a row-uniform global background and a non-uniform residual. The residual shows dataset-dependent effects. Overall, uniform full-range mixing provides a strong global spatial baseline, while the non-uniform attention residual is not consistently beneficial across datasets.

补充信息

↑