arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Pythia:面向多模态时间序列的基础世界模型

Pythia: Toward Foundation World Models for Multimodal Time Series

Xilin Dai, Hongzhou Chen, Yifan Hu, Yiding Liu, Zewei Dong, Jiang-Ming Yang

arXiv 2610.05240首次发表:更新:

发表机构

Ant International(蚂蚁国际)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Pythia提出联合嵌入预测架构的时间序列基础世界模型,通过上下文条件潜在动态和概率解码器解耦预训练与预测,在MUSE上显著降低误差,验证多模态信息互补贡献。

AI 中文摘要

时间序列基础模型提供了一种跨异构领域进行预测的统一方法。文本上下文和辅助观测为时间动态提供了互补信息,然而可复用的多模态预测表示仍未得到充分探索。我们引入了Pythia,一个基础世界模型,它通过联合嵌入预测架构学习跨数据集的上下文条件潜在动态。一个停止梯度的数值参考引导对预测未来状态的上下文修正。随后,一个独立的概率解码器适应冻结的预测表示和观测历史,将世界模型预训练与观测空间预测解耦。在MUSE上,Pythia-Tiny的归一化平均绝对缩放误差(MASE)和加权分位数损失(WSQL)分别为0.6879和0.4269,相对于已发布的MUSE排行榜中评估的最强模型,误差分别降低了6.26%和5.00%。通过一系列受控实验,我们研究了如何通过共享预训练设计时间序列世界模型,以及联合嵌入预测学习如何整合多模态信息。结果支持将预测表示学习与概率读出分离,并显示了实体描述、事件和协变量的互补贡献。

英文摘要

Time-series foundation models offer a unified approach to forecasting across heterogeneous domains. Textual context and auxiliary observations provide complementary information about temporal dynamics, yet reusable multimodal predictive representations remain underexplored. We introduce Pythia, a foundation world model that learns context-conditioned latent dynamics across datasets through a joint-embedding predictive architecture. A stop-gradient numerical reference guides contextual corrections to predicted future states. A separate probabilistic decoder then adapts to the frozen predictive representation and observed history, decoupling world-model pretraining from observation-space forecasting. On MUSE, Pythia-Tiny's normalized mean absolute scaled error (MASE) and weighted sum quantile loss (WSQL) are 0.6879 and 0.4269, reducing errors by 6.26% and 5.00% relative to the strongest model evaluated in the published MUSE leaderboard. Through a series of controlled experiments, we investigate how to design a time-series world model through shared pretraining and how joint-embedding predictive learning can incorporate multimodal information. The results support separating predictive representation learning from probabilistic readout and show complementary contributions from entity descriptions, events, and covariates.

CommentsTechnical Report

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑