发表机构
Carnegie Mellon University; Amazon; Nixtla(卡内基梅隆大学; 亚马逊; Nixtla)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出尺度不变训练(ScaleIn),通过修正尺度污染训练(ScaleCon)使时间序列基础模型对序列尺度不变,在24个架构-基准比较中平均降低MASE 18.8%和21.9%,且仅需一行代码改动。
AI 中文摘要
时间序列基础模型(TSFMs)在涵盖多种形态和领域的大型时间序列数据集集合上进行训练。这种设置使模型暴露于尺度(即其值的典型幅度)可能差异很大的序列。诸如可逆实例归一化(ReVIN)之类的仿射缩放方法对模型输入进行缩放,并在计算损失之前反转变换。我们证明,这种反转使得每个序列的梯度相对于缩放目标上的损失乘以$b^p$,其中$b$是缩放分母(例如标准差),$p$是损失次数。我们称之为尺度污染训练(ScaleCon),因为每个序列的尺度因此成为重要性权重,导致高尺度序列主导训练。对于任何尺度等变缩放器和齐次度为$p$的残差损失,包括MSE、MAE和分位数损失,我们证明在缩放目标上计算损失使得每个小批量梯度以及因此整个优化轨迹对训练序列的任意独立重新缩放不变,从而产生尺度不变训练(ScaleIn)。值得注意的是,现有的TSFM同时使用这两种目标,既没有一致的报告,也没有关于如何计算训练损失的通用约定。我们在合成和真实数据的受控研究中隔离了ScaleCon引起的收敛差异及其在ScaleIn下的修正。在四种TSFM架构的预训练中,ScaleIn在所有24个架构-基准比较中降低了MASE,在GIFT-Eval上平均降低18.8%,在M-竞赛上平均降低21.9%。这些增益扩展到监督神经预测,在20个匹配设置中的16个中降低了MASE。大多数现有的时间序列预测流程可以通过一行代码更改采用ScaleIn。
英文摘要
Time series foundation models (TSFMs) are trained on large collections of time series datasets that span various morphologies and domains. This setting exposes models to series whose scales -- typical magnitudes of their values -- can differ substantially. Affine scaling methods such as Reversible Instance Normalization (ReVIN) scale model inputs and reverse the transform before computing the loss. We show that this inversion multiplies each series' gradient by $b^p$ relative to loss on scaled targets, where $b$ is the scaling denominator (e.g., standard deviation) and $p$ is the loss degree. We call this scale-contaminated training (ScaleCon), because the scale of each series consequently becomes an importance weight, causing high-scale series to dominate training. For any scale-equivariant scaler and residual loss that is homogeneous of degree $p$, including MSE, MAE, and Quantile Loss, we prove that computing loss on scaled targets makes every mini-batch gradient and, consequently, the full optimization trajectory invariant to arbitrary independent rescaling of the training series, yielding scale-invariant training (ScaleIn). Notably, existing TSFMs use both objectives, with neither consistent reporting nor a common convention on how to compute training loss. We isolate the convergence disparity induced by ScaleCon and its correction under ScaleIn in controlled studies on synthetic and real data. In pretraining across four TSFM architectures, ScaleIn lowers MASE in all 24 architecture-benchmark comparisons, with average reductions across TSFMs of 18.8% on GIFT-Eval and 21.9% on the M-competitions. The gains extend to supervised neural forecasting, where it lowers MASE in 16 of 20 matched settings. Most existing time series forecasting pipelines can adopt ScaleIn with a one-line code change.