发表机构
Layer6 AI(Layer6 AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文通过SimpleTimeBench诊断套件发现,主流时间序列基础模型在基本时间模式及外生协变量利用上存在零样本盲区,微调反而损害泛化能力,揭示预训练规模与基本推理间的差距。
AI 中文摘要
尽管时间序列基础模型(TSFMs)在广泛基准测试中取得了成功,但它们在内在化基本时间逻辑方面的能力,尤其是在有外生协变量支持的场景中,仍未得到充分检验。我们引入了SimpleTimeBench,一个用于单调趋势、周期信号和领先指标协变量等基本要素的诊断性单变量和多变量“单元测试”套件,在这些场景中,近乎完美的预测应当是轻而易举的。令人惊讶的是,著名的多变量TSFMs(Chronos-2、Moirai和Toto)在处理这些输入时经常产生次优的零样本预测。虽然对Chronos-2进行微调能改善其在特定任务上的表现,但我们表明,这种适应会降低其在其他基本模式上的性能,而非增强其可泛化的基础能力。这揭示了预训练规模与基本时间推理之间的差距,表明当前的TSFMs可能缺乏捕捉简单可预测函数所需的归纳偏置。我们进一步证明,这些失败并非仅仅是合成数据中的奇特现象:它们在真实世界的传感器预测中持续存在,其中TSFMs始终未能充分利用观测协变量中可用的领先指标。这种无法捕捉简单关系的能力限制了当前多变量模型的实际效用和可靠性。
英文摘要
Despite the success of Time Series Foundation Models (TSFMs) on broad benchmarks, their ability to internalize basic temporal logic, especially in settings supported by exogenous covariates, remains under-examined. We introduce SimpleTimeBench, a diagnostic univariate and multivariate "unit test" suite for primitives such as monotonic trends, periodic signals and leading indicator covariates, scenarios where near-perfect forecasts should be trivial. Surprisingly, prominent multivariate TSFMs (Chronos-2, Moirai and Toto) frequently produce suboptimal zero-shot forecasts for these inputs. While fine-tuning Chronos-2 improves its behaviour on specific tasks, we show that this adaptation degrades performance on other fundamental patterns rather than enhancing its generalizable foundational capabilities. This reveals a gap between pre-training scale and basic temporal reasoning, suggesting that current TSFMs could potentially lack the inductive biases needed to capture simple predictable functions. We further demonstrate that these failures are not merely synthetic curiosities: they persist in real-world sensor forecasting, where TSFMs consistently underutilize leading indicators available in observed covariates. This inability to capture simple relationships limits the practical utility and reliability of current multivariate models.