发表机构
Alexandria University(亚历山大大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过模拟数据比较谷歌TimesFM-3与经典时间序列模型,发现两者各有优劣,TimesFM-3在短视野稳健但长视野不足,经典方法在季节数据上更优。
AI 中文摘要
时间序列基础模型(如谷歌的TimesFM-3)能够预测它们从未训练过的序列,但它们是否应取代ARIMA和指数平滑等经典方法,在公开基准上难以定论,因为很少有基准数据集完全不在每个模型的预训练语料中。这项探索性研究在从九个已知过程生成的数据上比较了TimesFM-3与经典方法,这些数据模型不可能见过。在7200个序列和预注册的12步预测视野协议下,两个系列均未占优。TimesFM-3的误差在任一场景中最多为最佳方法的1.31倍,而每个经典方法在某个场景中至少为2.2倍;最大的经典方法失败发生在仅有两个季节周期时,自动方法退化为非季节模型。在四个或更多周期时,自动ARIMA在季节ARIMA数据上比TimesFM-3准确21-24%。在五个基础模型中,这种稳健性为TimesFM-3所特有,且仅在短视野成立:在第25至48步,所测试的三个基础模型在水平保持平坦的饱和过程上预测下降,而自动ARIMA在其他地方具有较小的最坏情况。在间歇性需求上,TimesFM-3相对于基于历史构建的简单基准没有优势。随机参数、重尾噪声和异常值在定性上不改变这些模式;在官方测试期的1000个M4月度序列上,TimesFM-3与M4排名第五至第七的参赛者持平,落后于前四名,在模型文档化训练数据之后观测的101个宏观经济序列上,它与经典方法统计上持平。该研究刻画了黑盒模型的行为,而非其原因,其发现以所考察的过程和视野为条件。
英文摘要
Time-series foundation models such as Google's TimesFM-3 forecast series they were never trained on, but whether they should replace classical methods such as ARIMA and exponential smoothing is hard to settle on public benchmarks, because few benchmark datasets are absent from every model's pre-training corpus. This exploratory study compares TimesFM-3 with classical methods on data generated from nine known processes, which the model cannot have seen. Across 7200 series and a pre-registered protocol with a 12-step horizon, neither family dominates. TimesFM-3's error is never more than 1.31 times that of the best method in a scenario, whereas every classical method's is at least 2.2 times somewhere; the largest classical failures occur with only two seasonal cycles, where the automatic methods fall back to non-seasonal models. With four or more cycles automatic ARIMA is 21-24% more accurate than TimesFM-3 on seasonal ARIMA data. Among five foundation models this robustness is specific to TimesFM-3, and it holds only at short horizons: at steps 25 to 48 the three foundation models tested there forecast a decline on a saturating process whose level stays flat, and automatic ARIMA has the smaller worst case elsewhere. On intermittent demand TimesFM-3 shows no advantage over simple benchmarks built from the history. Randomised parameters, heavy-tailed noise and outliers leave these patterns qualitatively unchanged; on the official test period of 1000 M4 monthly series TimesFM-3 is level with the fifth- to seventh-ranked M4 entries, behind the four best, and on 101 macroeconomic series observed after the models' documented training data it is statistically level with the classical methods. The study characterises the behaviour of a black-box model, not the reasons for it, and its findings are conditional on the processes and horizons examined.
Comments40 pages (23 main text and a 17-page Supporting Information), 9 figures, 26 tables. Code, data, protocol and results: https://doi.org/10.5281/zenodo.22681072