发表机构
East China Normal University(华东师范大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Aurora-X是一个十亿规模的时间序列基础模型,通过渐进式课程学习、统一架构和模式引导的专家混合,实现了跨变量建模与概率预测,在多个基准上达到最先进性能。
AI 中文摘要
时间序列基础模型(TSFMs)能够实现跨域预测,但其作为通用预测器的发展仍受限于未充分探索的训练潜力和有限的架构通用性。为应对这些挑战,我们提出了Aurora-X,一个十亿规模的时间序列基础模型,具有渐进式课程学习和统一架构。我们首先使用通道独立预训练来学习时间模式,然后在中期训练中引入跨变量依赖、变化的上下文和预测长度,以及未来协变量(如果可用)。可变分辨率后训练进一步支持推理时每个令牌可调整的时间跨度。在固定模型权重下,这支持在固定令牌预算下处理更长的历史,或对相同历史使用更少的令牌,从而实现测试时扩展。凭借通用架构,Aurora-X支持跨变量建模、协变量条件化以及未来补丁的并行解码以进行概率预测。这些由一个新颖的模式引导的专家混合模型支持,该模型通过稀疏激活扩展模型容量,并利用浅层补丁相似性来约束深层路由,引导专家在异构时间序列中的专业化。此外,我们提出了一个隐式分位数网络头,可预测任意分位数以表征预测分布,增强概率预测的灵活性。在GIFT-Eval、TIME、FEV-Bench、TFB和DAG-Bench上的综合实验表明,与预训练的时间序列基础模型和任务特定的监督模型相比,我们的模型达到了最先进的预测性能。
英文摘要
Time series foundation models (TSFMs) enable cross-domain forecasting, but their development as general-purpose forecasters remains constrained by underexplored training potential and limited architectural versatility. To address these challenges, we introduce Aurora-X, a billion-scale TSFM with a progressive curriculum and a unified architecture. We first use channel-independent pretraining to learn temporal patterns, then introduce cross-variable dependencies, varied context and horizon lengths, and future covariates if available during midtraining. Variable-resolution post-training further enables an adjustable temporal span per token at inference. With fixed model weights, this supports longer histories under a fixed token budget or fewer tokens for the same history, enabling test-time scaling. With a versatile architecture, Aurora-X supports cross-variable modeling, covariate conditioning, and parallel decoding of future patches for probabilistic forecasting. These are supported by a novel pattern-guided mixture-of-experts that expands model capacity through sparse activation and uses shallow patch similarities to constrain deep-layer routing, guiding expert specialization across heterogeneous time series. Furthermore, we propose an implicit quantile network head that predicts arbitrary quantiles to characterize predictive distributions, enhancing probabilistic forecasting flexibility. Comprehensive experiments on GIFT-Eval, TIME, FEV-Bench, TFB, and DAG-Bench demonstrate state-of-the-art forecasting performance against pretrained TSFMs and task-specific supervised models.