发表机构
Ant International(蚂蚁国际)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Pythia提出联合嵌入预测架构的时间序列基础世界模型,通过上下文条件潜在动态和概率解码器解耦预训练与预测,在MUSE上显著降低误差,验证多模态信息互补贡献。
AI 中文摘要
时间序列基础模型提供了一种跨异构领域进行预测的统一方法。文本上下文和辅助观测为时间动态提供了互补信息,然而可复用的多模态预测表示仍未得到充分探索。我们引入了Pythia,一个基础世界模型,它通过联合嵌入预测架构学习跨数据集的上下文条件潜在动态。一个停止梯度的数值参考引导对预测未来状态的上下文修正。随后,一个独立的概率解码器适应冻结的预测表示和观测历史,将世界模型预训练与观测空间预测解耦。在MUSE上,Pythia-Tiny的归一化平均绝对缩放误差(MASE)和加权分位数损失(WSQL)分别为0.6879和0.4269,相对于已发布的MUSE排行榜中评估的最强模型,误差分别降低了6.26%和5.00%。通过一系列受控实验,我们研究了如何通过共享预训练设计时间序列世界模型,以及联合嵌入预测学习如何整合多模态信息。结果支持将预测表示学习与概率读出分离,并显示了实体描述、事件和协变量的互补贡献。
英文摘要
Time-series foundation models offer a unified approach to forecasting across heterogeneous domains. Textual context and auxiliary observations provide complementary information about temporal dynamics, yet reusable multimodal predictive representations remain underexplored. We introduce Pythia, a foundation world model that learns context-conditioned latent dynamics across datasets through a joint-embedding predictive architecture. A stop-gradient numerical reference guides contextual corrections to predicted future states. A separate probabilistic decoder then adapts to the frozen predictive representation and observed history, decoupling world-model pretraining from observation-space forecasting. On MUSE, Pythia-Tiny's normalized mean absolute scaled error (MASE) and weighted sum quantile loss (WSQL) are 0.6879 and 0.4269, reducing errors by 6.26% and 5.00% relative to the strongest model evaluated in the published MUSE leaderboard. Through a series of controlled experiments, we investigate how to design a time-series world model through shared pretraining and how joint-embedding predictive learning can incorporate multimodal information. The results support separating predictive representation learning from probabilistic readout and show complementary contributions from entity descriptions, events, and covariates.
CommentsTechnical Report