arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当注意力无法解释峰值:基于注意力的时间序列预测中的时间参考与预测输出

When Attention Does Not Explain the Peak: Temporal Reference vs. Forecast Output in Attention-Based Time-Series Forecasting

Yuji Akamatsu, Takao Yamanaka

arXiv 2610.07080首次发表:更新:

发表机构

Sophia University(上智大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过巴拿马负荷数据集的31个窗口实验,发现时域级交叉注意力的峰值选择性与预测峰值无关,其作用更可能是整合未来外生信息的时间参考,而非解释预测行为。

AI 中文摘要

注意力图常被解释为预测模型在做预测时所使用信息的证据。在我们的负荷预测模型中,历史需求的一个CLS表示通过交叉注意力查询24个未来外生时域标记,这引发了一种时间解释,即高注意力时域可能看似解释了预测峰值的时序。我们使用时域级注意力描述符 $\Psi_{\mathrm{out}}$ 来检验这一解释。在巴拿马负荷数据集的31个按日对齐窗口中,预测的中位峰值时间误差为0小时,精确匹配率为51.6%,而 $\Psi_{\mathrm{out}}$ 的argmax中位误差为5小时,精确匹配率为0%。在31个窗口中,有27个窗口的预测峰值更接近观测峰值。这种分离并非仅仅是argmax的伪影:在±1小时内,注意力在观测峰值、预测峰值和周朴素峰值周围的浓度分别仅为均匀基线的1.16倍、1.11倍和1.14倍,表明注意力集中度弱且非选择性。然而,注意力分布是有结构的,跨窗口一致性为0.83。将12个未来天气特征替换为其训练集均值会使分布近乎均匀,这表明注意力对未来天气变化敏感,而非仅对固定时域位置敏感。这种分离也在三个随机种子运行中重现。这些结果表明,结构化、对输入敏感且可复现的时域级交叉注意力未必能为预测行为提供有效的峰值选择性解释。观测到的行为反而与一种内部时域参考作用一致,即用于整合未来外生信息,尽管这一功能作用尚未被因果性地确立。

英文摘要

Attention maps are often interpreted as evidence of what a forecasting model uses when making predictions. In our load-forecasting model, a CLS representation of historical demand queries 24 future exogenous horizon tokens through cross-attention, inviting a temporal interpretation in which highly attended horizons may appear to explain forecast peak timing. We test this interpretation using a horizon-level attention descriptor, $Ψ_{\mathrm{out}}$. Across 31 day-aligned windows of the Panama load dataset, the forecast achieves a median peak-time error of 0 h and a 51.6% exact-match rate, whereas the argmax of $Ψ_{\mathrm{out}}$ has a median error of 5 h and 0% exact match. The forecast peak is closer to the observed peak in 27 of 31 windows. This dissociation is not merely an argmax artifact: within $\pm1$ h, attention reaches only $1.16\times$, $1.11\times$, and $1.14\times$ the uniform baseline around observed, predicted, and weekly-naive peaks, respectively, indicating weak and non-selective concentration. Yet the attention profile is structured, with cross-window consistency of 0.83. Replacing 12 future weather features with their training-set means makes the profile nearly uniform, showing sensitivity to future weather variation rather than fixed horizon position alone. The dissociation is also reproduced across three random-seed runs. These results show that structured, input-sensitive, and reproducible horizon-level cross-attention need not provide a valid peak-selective explanation of forecast behavior. The observed behavior is instead consistent with an internal horizon-reference role for integrating future exogenous information, although this functional role is not causally established.

CommentsNeurIPS 2026: Accepted to the TAE (Trust-AI-Eval) Workshop

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑