发表机构
National School of Computer Sciences (ENSI), University of Manouba, Manouba, Tunisia(突尼斯曼努巴大学计算机科学国家学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究预训练生成式基础模型,通过功率控制研究、受控合成实验等得出结果,表明其优势与趋势强度有关,趋势强度可作为先验指标指导基础模型部署,还记录了校准不足。
AI 中文摘要
预训练生成式基础模型将预测视为从学习到的预测分布进行条件生成,并能零样本预测未见序列。我们得出三个结果,将其已报道的成功转化为可操作的、有机制的理解。首先,在对跨越STL趋势强度全范围(\(F_T\in[0.17,1.00]\))的19个数据集的36个序列进行的1728次滚动原点预测的功率控制研究中,零样本Chronos模型显著优于四个强大的经典基线。其次,通过一个已知趋势生成过程的受控合成实验揭示了获胜原因及并非源于更好的趋势外推。最后,由于优势是收缩效应,仅从趋势强度就能预测,趋势强度是部署基础模型的实用先验指标,我们还记录了校准不足。
英文摘要
Pretrained generative foundation models cast forecasting as conditional generation from a learned predictive distribution and forecast unseen series zero-shot. We establish three results that turn their reported success into an actionable, mechanistic understanding. First (a positive benchmark result): on a power-controlled study of 1728 rolling-origin forecasts over 36 series from 19 datasets spanning the full range of STL trend strength (F_T in [0.17, 1.00]), a zero-shot Chronos model significantly outperforms four strong classical baselines -- drift, seasonal-naive, Theta, and additive Holt-Winters/ETS -- with the best mean MASE (1.187 vs. Theta 1.337, ETS 1.656; Friedman chi^2 = 46.08, p = 8.75e-09; Holm-corrected Wilcoxon p <= 0.015 against every baseline; a Nemenyi critical difference separating it from the classical pack). Second (a novel, quantified mechanism): a controlled synthetic experiment with a known trend-generating process shows why -- and reveals that the win does not come from better trend extrapolation. When the true trend is linear, damped, or exponential, additive ETS tracks the slope (slope-tracking ratio 1.02, 1.34, 0.98) whereas Chronos systematically under-extrapolates, behaving as a trend-shrinkage estimator (ratio 0.80, 0.49, 0.36). Third (an actionable selection rule): because the advantage is a shrinkage effect, it is predictable from trend strength alone -- the generative model wins 78% of low-trend series but only 44% of high-trend ones, and its edge over ETS is significant on the low-trend stratum (0.982 vs. 1.671, p = 0.002) yet a tie on the high-trend stratum (p = 0.18). Trend strength, computable before forecasting from the training context alone, is therefore a practical a-priori indicator of when to deploy a foundation model. We additionally document a calibration shortfall (80% intervals cover 0.77).
Comments12 pages, 6 figures. Code and per-forecast data released