发表机构
Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种静态数据选择框架,通过参考预测器评分并保留中间区间,在减少预训练窗口的同时提升时间序列基础模型的性能,且小型参考模型可有效为大型模型选择数据。
AI 中文摘要
时间序列基础模型(TSFMs)在包含数十亿观测值的异构集合上进行预训练,然而其训练窗口通常被采样时并未评估它们是否提供有用的学习信号。我们引入了一个静态数据选择框架,该框架使用参考预测器对每个窗口进行评分,并在每个源数据集中保留一个中间区间。具体而言,我们通过展示在局部雅可比条件下归一化平方损失控制每个样本的梯度范数,将预测损失与优化难度联系起来。然后我们定义了一个参考损失分数,并应用数据集分层选择以保留样本的多样性。在各种TSFM架构中,我们的方法以绝对边际优于随机选择,并且通过保留更少的候选预训练窗口,甚至相对于全数据预训练提高了相对MASE和CRPS。进一步的分析显示了跨尺度和跨架构的强分数相关性,表明只要参考模型和目标模型具有兼容的难度排序,一个小型参考模型通常可以为更大的目标模型选择数据。
英文摘要
Time series foundation models (TSFMs) are pretrained on heterogeneous collections containing billions of observations, yet their training windows are typically sampled without estimating whether they provide useful learning signal. We introduce a static data-selection framework that scores each window with a reference forecaster and retains an intermediate interval within every source dataset. Specifically, we connect forecasting loss to optimization difficulty by showing that normalized squared loss controls the per-sample gradient norm under a local Jacobian condition. We then define a reference loss score and apply dataset-stratified selection to preserve the diversity of samples. Across various TSFM architectures, our method outperforms random selection by an absolute margin and even improves both relative MASE and CRPS over full-data pretraining by retaining fewer candidate pretraining windows. Further analyses show strong cross-scale and cross-architecture score correlations, indicating that a small reference model can often select data for larger targets, provided that the reference and target share compatible difficulty orderings.
Comments15 pages, 5 figures