发表机构
Meta AI(Meta AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出合成数据蒸馏(SDD)方法,通过比较模型输出与条件预测分布来训练时间序列基础模型,减少梯度协方差,加速收敛并降低训练迭代次数。
AI 中文摘要
时间序列基础模型(TSFMs)越来越多地在合成生成的时间序列轨迹上进行预训练,其中数据生成过程是已知的。当前的预训练方法基于损失目标,将TSFM的输出与每条轨迹的实现未来值进行比较。我们转而提出损失目标,将TSFM的输出与每条轨迹的条件预测分布进行比较,这一过程我们称之为合成数据蒸馏(SDD)。SDD对应于训练目标的Rao-Blackwell化,即它保持随机梯度的期望不变,同时在Loewner偏序下可证明地减少随机梯度的协方差。我们在参数规模从4M到2.5B的TSFM模型家族上实证验证了SDD,观察到在每个模型规模下验证损失的收敛速度更快:在高斯过程数据上,SDD达到或优于现状损失,同时所需训练迭代次数减少10%-40%。
英文摘要
Time series foundation models (TSFMs) are increasingly pre-trained on synthetically generated time series trajectories, where the data generating process is known. Current pre-training recipes are based on loss objectives which compare TSFM outputs to realized future values of each trajectory. We instead propose loss objectives which compare TSFM outputs to the conditional forecast distribution of each trajectory, a procedure we call synthetic data distillation (SDD). SDD corresponds to a Rao-Blackwellization of the training objective, in that it leaves the expectation of stochastic gradients unchanged while provably reducing the covariance of the stochastic gradient under the Loewner partial ordering. We empirically validate SDD on a TSFM model family of sizes from $4$M to $2.5$B parameters, and observe faster convergence of validation loss at every model size: on Gaussian Process data, SDD attains or improves upon the Status Quo loss whilst requiring $10\%-40\%$ less training iterations.
Comments13 pages, 3 figures