发表机构
LIPADE, Université Paris Cité; École Polytechnique; Centre de Recherche en Informatique, Mines Paris, PSL University; IRISA, Université Rennes 2, Inria(巴黎西岱大学LIPADE实验室; 巴黎综合理工学院; 巴黎高等矿业学院计算机研究中心,巴黎文理研究大学; 雷恩第二大学IRISA实验室,法国国家信息与自动化研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FlowTSFM通过共享参数的循环Transformer块和分位数流目标,将编码器深度转化为传输过程,以38.8M参数在GIFT-Eval和TIME上接近更强基线,并显著提升预测轨迹的结构化程度。
AI 中文摘要
基于编码器的时间序列基础模型(TSFM)通常依赖于深度堆叠的独立参数化Transformer层,其中仅最终预测受到监督,中间表示没有明确的预测作用。我们引入了FlowTSFM,一种编码器架构,将深度解释为循环传输过程:单个Transformer块以共享参数迭代应用,而分位数流目标沿从先验分布到最终预测的规定轨迹监督中间状态。该目标结合了钉球预测损失与路径级位置匹配。仅用38.8M参数,FlowTSFM在GIFT-Eval和TIME上取得了有竞争力的性能,与更强的基线相比,MASE保持在1.8-4.6%以内,同时使用的参数比12层Chronos-2模型(119.5M)少约3倍。除了准确性之外,我们引入了CosMean,一种无标度诊断指标,用于衡量循环更新是否一致地对齐到最终预测。在匹配的中间状态探测协议下,FlowTSFM的CosMean得分为0.919,而Chronos-2为0.350,这表明循环参数共享结合路径监督与更结构化的预测轨迹以及有利的准确性-效率权衡相关联。
英文摘要
Encoder-based time series foundation models (TSFMs) typically rely on deep stacks of independently parameterized Transformer layers, where only the final forecast is supervised and intermediate representations have no explicit predictive role. We introduce FlowTSFM, an encoder architecture that interprets depth as a recurrent transport process: a single Transformer block is iteratively applied with shared parameters, while a quantile-flow objective supervises intermediate states along a prescribed trajectory from a prior distribution toward the final forecast. The objective combines pinball forecasting loss with path-level position matching. With only 38.8M parameters, FlowTSFM achieves competitive performance on GIFT-Eval and TIME, remaining within 1.8-4.6% MASE of stronger baselines while using approximately $3\times$ fewer parameters than a 12-layer Chronos-2 model (119.5M). Beyond accuracy, we introduce CosMean, a scale-free diagnostic measuring whether recurrent updates consistently align toward the final prediction. Under a matched intermediate-state probing protocol, FlowTSFM achieves a CosMean score of 0.919 compared with 0.350 for Chronos-2, suggesting that recurrent parameter sharing combined with path supervision is associated with substantially more structured predictive trajectories at a favorable accuracy-efficiency trade-off.