评估零样本时间序列基础模型的准确性与概率可靠性
Evaluating Accuracy and Probabilistic Reliability of Zero-Shot Time Series Foundation Models
查看机构详情
- University of Nicosia(尼科西亚大学)
- University of Cyprus(塞浦路斯大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文基准评估六种零样本时间序列基础模型,发现其优于统计与监督方法,但存在点准确性与概率校准的权衡,xLSTM校准稳健,补丁Transformer长程校准欠佳。
中文摘要 AI 辅助
时间序列基础模型(TSFMs)承诺通过消除任务特定训练来实现向零样本预测的范式转变。然而,现有工作往往忽视了预测准确性与概率校准之间的权衡。本文对六种TSFM在能源、交通和金融数据集上的表现进行了基准研究。我们将它们的性能与统计基线和监督深度学习模型进行了对比。研究揭示,尽管TSFM优于统计方法和监督模型,但它们面临点准确性与概率可靠性之间的根本性权衡。具体而言,xLSTM架构在不同预测范围上提供了稳健的概率校准。相比之下,基于补丁的Transformer在长预测范围上具有竞争力的准确性但面临校准问题,而基于Transformer的模型则表现出用于最优零样本推理的上下文饱和点。这些发现为在实际部署中平衡泛化与不确定性量化提供了基于证据的指导。
英文摘要
Time Series Foundation Models (TSFMs) promise a paradigm shift toward zero-shot forecasting by eliminating task-specific training. However, existing works often overlook trade-offs between predictive accuracy and probabilistic calibration. This paper presents a benchmark study of six TSFMs evaluated on energy, traffic, and financial datasets. We contrast their performance against statistical baselines and a supervised DL model. The study reveals that while TSFMs outperform statistical methods and supervised models, they are subject to a fundamental trade-off between point accuracy and probabilistic reliability. Specifically, xLSTM architectures provide robust probabilistic calibration across horizons. In contrast, patch-based transformers offer competitive accuracy but face calibration issues at long horizons, while transformer-based models exhibit context saturation points for optimal zero-shot reasoning. These findings offer evidence-based guidance for balancing generalization and uncertainty quantification in real-world deployments.