何时而非多少:评估时间序列基础模型在稀疏事件上的表现
When, Not How Much: Evaluating Time-Series Foundation Models on Sparse Events
浏览论文内容
中文总结 AI 辅助
本研究评估时间序列基础模型在稀疏事件排序任务上的表现,发现发布的点预测提升有限,而冻结模型上的轻量级事件头能更有效地排序事件,并超越部分基线。
中文摘要 AI 辅助
预训练的时间序列基础模型(TSFMs)被评估为未来值的预测器,然而对于稀疏序列,许多决策仅取决于哪些未来时期包含活动。标准基准测试并未评估这一点。在五个稀疏数据集上,我们对预测窗口内同时包含事件和零的位置进行排序。12个TSFM发布的点预测相对于无训练参考,在机会校正平均精度上最多提升0.031,而在机会校正AUC上,中位数TSFM在每个数据集上都低于这些参考。在事件监督下,六个冻结骨干网络的线性探针在29个骨干-数据集对(共30个)中优于其骨干的点预测。对预测的分位数取平均而非中位数,改善了对大多数预测中位数的TSFM的排序,在两个数据集上,最强的此类输出可与探针相媲美。探针相对于原始上下文学习者的优势取决于数据集,在同一探针下,预训练特征在六个骨干中的五个上优于随机初始化的特征。对于稀疏事件排序,发布的点预测相对于简单参考几乎没有增加价值,而冻结TSFM上的轻量级事件头比这些预测更好地对事件进行排序,其中最好的在五个数据集中的三个上超过了在原始上下文上训练的梯度提升树。更广泛地说,在价值预测之外的任务上评估预训练预测器,需要并排报告其输出、其表示的有监督探针以及原始上下文和随机化对照,因为每个都支持不同的结论。
英文摘要
Pretrained time-series foundation models (TSFMs) are evaluated as forecasters of future values, yet for sparse series many decisions depend only on which future periods contain activity. Standard benchmarks do not assess this. On five sparse datasets, we rank positions within forecast windows that contain both events and zeros. The released point forecasts of 12 TSFMs improve chance-corrected average precision over training-free references by at most 0.031, and in chance-corrected AUC the median TSFM falls below them on every dataset. With event supervision, linear probes of six frozen backbones improve on their backbone's point forecast in 29 of 30 backbone--dataset pairs. Averaging the predicted quantiles instead of taking their median improves the ranking of most TSFMs that forecast the median, and on two datasets the strongest such outputs rival the probes. The probes' advantage over raw-context learners depends on the dataset, and under the same probe, pretrained features outperform randomly initialized ones for five of six backbones. For sparse-event ranking, released point forecasts thus add little over simple references, whereas lightweight event heads on frozen TSFMs rank events better than these forecasts, and the best of them exceed gradient-boosted trees trained on the raw context on three of the five datasets. More broadly, assessing pretrained forecasters on tasks beyond value forecasting requires reporting their outputs, supervised probes of their representations, and raw-context and randomized controls side by side, since each supports a different conclusion.
发表机构
- ETH Zurich(苏黎世联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。