arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22820cs.LG

超越平均误差:面向时间序列预测的预言机知情压力测试

Beyond Average Error through Oracle-Informed Stress Tests for Time-Series Forecasting

Xu Lin, Runheng Zuo, Shengxuan Xu, Qitai Tan, Hongyu Lin, Xiao-Ping Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出预言机知情压力测试,将预测误差分解为环境风险与预言机距离,并在24个模型上验证,发现性能下降主要由环境风险驱动,且部分发现不可复现。

中文摘要 AI 辅助

平均平方误差无法揭示预测性能下降是因为未来变得难以预测,还是因为预测值偏离条件均值更远。我们引入了配对的、机制控制的压力测试,利用评估模型无法获得的、以起点为条件的预测预言机,将每个提前期预期平方误差的变化分解为环境风险和预测-预言机距离。三种端到端控制具有已知的归因。具体而言,零控制、仅环境控制和信息缺口控制验证了流程将变化归因于正确的组成部分。然后,我们将该基准应用于24个可部署的预测器。在频繁切换下,14种方法的实际均方误差更高但预言机距离更低;在异常值方差反馈下,19种方法的均方误差更高但尺度标准化均方误差更低。短提前期和长提前期的压力响应排名具有0.624的斯皮尔曼相关性,揭示了显著的依赖于时间跨度的重排序。接着,我们研究了多变量关系变化。在六种模型和三种耦合强度下,预言机距离仅占分解预期风险增加的0.7%-3.9%,并且在八通道系统以及对Ring、Block和Hub关系的匹配难度审计中,环境风险占主导地位。最后,对独立数据生成过程(DGP)实现的预先指定对比显示,几个视觉上引人注目的发现模式,包括趋势累积和假设的切换反转,并未复现。因此,该基准将分量诊断与留出稳定性审计相结合。它补充了真实数据分布外评估,后者在无法获得精确预言机归因时,衡量现实变化下的性能。

英文摘要

Average squared error cannot reveal whether forecasting performance degrades because the future becomes less predictable or because forecasts move farther from the conditional mean. We introduce paired, mechanism-controlled stress tests that decompose changes in expected squared error at each lead time into environmental risk and forecast-oracle distance, using an origin-conditioned predictive oracle unavailable to the evaluated models. Three end-to-end controls have known attribution. Specifically, the null, environmental-only, and information-gap controls verify that the pipeline assigns changes to the correct component. We then apply the benchmark to 24 deployable forecasters. Under frequent switching, 14 methods have higher realized MSE but lower oracle distance; under outlier-variance feedback, 19 have higher MSE but lower scale-standardized MSE. Short- and long-lead stress-response rankings have Spearman correlation 0.624, revealing substantial horizon-dependent reordering. We then study multivariate relation shifts. Across six models and three coupling severities, oracle distance accounts for only 0.7-3.9% of the decomposed expected-risk increase, and environmental-risk majority persists in an eight-channel system and a matched-difficulty audit of Ring, Block, and Hub relations. Finally, prespecified contrasts on independent data-generating process (DGP) realizations show that several visually compelling discovery profiles, including trend accumulation and the hypothesized switching reversal, do not replicate. The benchmark thus combines component-wise diagnosis with a held-out stability audit. It complements real-data out-of-distribution evaluation, which measures performance under realistic shifts when exact oracle attribution is unavailable.

发表机构

  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑