发表机构
J.P. Morgan AI Research; Imperial College London(摩根大通人工智能研究院; 伦敦帝国理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出时间预测多重性框架,揭示同等准确的时间序列模型在完整预测轨迹上存在显著分歧,且与逐时域分歧无关,对下游决策有重要影响。
AI 中文摘要
具有几乎相同预测性能的模型可能产生截然不同的预测,这一现象被称为预测多重性。先前的工作主要是在单个标量输出的层面上研究这一现象。然而,在时间序列预测中,跨多个预测时域的预测共同定义了一条轨迹,而逐时域的比较可能掩盖预测行为中的重要差异。为了解决这一问题,我们引入了时间预测多重性,这是一个描述具有几乎相同预测性能的模型之间在完整预测轨迹上存在分歧的框架。我们表明,仅约束预测性能仍然可以允许出现广泛多样的不同轨迹。我们进一步表明,在单个时域上约束多重性可以部分减少但无法消除轨迹层面的多重性。在11个数据集上对19种神经预测架构进行的实验证实,接近最优的模型在其产生的预测轨迹上可能表现出显著的变异性,并且轨迹层面的分歧与逐时域的分歧在很大程度上不相关。因此,我们的框架揭示了现有多重性研究中的一个空白:具有无法区分的预测性能的模型意味着根本不同的时间轨迹,并产生相应的下游影响。
英文摘要
Models with near-identical predictive performance can yield substantially different predictions, a phenomenon known as predictive multiplicity. Prior work has mostly studied this at the level of individual scalar outputs. In time-series forecasting, however, predictions across horizons jointly define a trajectory, and horizon-wise comparisons can hide important differences in predictive behavior. To address this problem, we introduce temporal predictive multiplicity, a framework that characterizes disagreement over complete forecast trajectories among models with near-identical predictive performance. We show that constraining predictive performance alone can still admit a broad range of different trajectories. We further show that constraining multiplicity at individual horizons partially reduces, but does not eliminate, trajectory-level multiplicity. Experiments with 19 neural forecasting architectures on 11 datasets confirm that near-optimal models can exhibit substantial variability in the forecast trajectories they produce, and trajectory-level disagreement is largely unrelated to horizon-wise disagreement. Our framework, therefore, exposes a gap in existing multiplicity studies: models with indistinguishable predictive performance imply fundamentally different temporal trajectories, with consequential downstream effects.