发表机构
The Hong Kong University of Science and Technology (Guangzhou); The University of Western Australia(香港科技大学(广州); 西澳大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对时间序列预测评估过窄的问题,提出基于场景的压力测试,每个测试实例包含语义场景、失败操作符和难度级别,使评估可解释、可归因且与部署相关。
AI 中文摘要
时间序列预测(TSF)日益驱动交通、能源、金融、医疗和基础设施领域的决策,然而当前的评估仍然过于狭窄:标准基准奖励较低的留出误差,而鲁棒性研究通常将失败简化为高斯噪声、随机掩蔽或有界对抗扰动。这掩盖了部署预测系统的真实失败模式。输入侧异常不仅仅是噪声更大的输入:它们通常反映结构化事件,这些事件改变时间动态、破坏跨变量依赖、引发制度转换,或从故障传感器传播到下游决策。这些语义性、因果性和系统级失败无法仅通过i.i.d.扰动忠实捕获。TSF基础模型的兴起使这一评估差距更加紧迫,因为未经审计的预训练语料库使留出泛化越来越不可靠。因此,我们倡导基于场景的压力测试。每个测试实例应包含历史输入和未来目标,以及语义场景、显式失败操作符和可测量的难度级别。这一转变使评估可解释、可归因、与部署相关且友好,使社区不仅能够询问哪个模型准确,还能询问它在什么条件下失败以及为什么失败。
英文摘要
Time series forecasting (TSF) increasingly drives decisions in transportation, energy, finance, healthcare, and infrastructure, yet current evaluation remains overly narrow: standard benchmarks reward low held-out error, while robustness studies typically reduce failure to Gaussian noise, random masking, or bounded adversarial perturbations. This obscures the real failure modes of deployed forecasting systems. Input-side anomalies are not merely noisier inputs: they often reflect structured events that alter temporal dynamics, break cross-variable dependencies, induce regime shifts, or propagate from faulty sensors to downstream decisions. These semantic, causal, and system-level failures cannot be faithfully captured by i.i.d. perturbations alone. The rise of TSF foundation models makes this evaluation gap more urgent, as unauditable pretraining corpora make held-out generalization increasingly unreliable. We therefore advocate scenario-grounded stress testing. Each test instance should include historical inputs and future targets, together with a semantic scenario, an explicit failure operator, and a measurable difficulty level. This shift makes evaluation interpretable, attributable, and deployment-relevant and friendly, enabling the community to ask not only which model is accurate, but under what conditions it fails and why.