发表机构
East China Normal University; Huawei Technologies Co., Ltd.(华东师范大学; 华为技术有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
QiYao-M提出角色感知的多模态时间序列基础模型,分别建模内源与外源模态,通过内源预测器、外源检索增强器及代理训练,在单/多模态基准上实现强预测性能。
AI 中文摘要
现有的多模态时间序列基础模型(TSFMs)通常通过大体共享的机制对异构模态进行建模,忽视了内源和外源模态在预测中的不同角色。在这项工作中,我们提出了QiYao-M,一种角色感知的多模态TSFM,它分别对这两种类型的模态进行建模。对于内源模态,为了捕捉它们如何随潜在的时间动态演化,我们引入了内源多模态预测器和内源多模态监督,以显式学习它们从历史到未来的演化。对于外源模态,为了在外源多模态预训练数据稀缺的情况下跨领域以及跨各种模态类型和数量进行泛化,我们提出了外源多模态检索增强器,它能够在不更新TSFM参数的情况下实现快速的下游适应。我们进一步引入了内源模态代理训练,以在没有外源多模态预训练数据的情况下训练该检索模块。在单模态和多模态基准上的广泛实验表明,在有无外源模态的场景中均表现出强大的预测性能。
英文摘要
Existing multimodal time series foundation models (TSFMs) typically model heterogeneous modalities through largely shared mechanisms, overlooking the distinct forecasting roles of endogenous and exogenous modalities. In this work, we propose QiYao-M, a role-aware multimodal TSFM that models the two types of modalities separately. For endogenous modalities, to capture how they evolve along with the underlying temporal dynamics, we introduce an Endo-Multimodal Predictor and Endo-Multimodal Supervision to explicitly learn their evolution from history to the future. For exogenous modalities, to generalize across domains and across various modality types and numbers under the scarcity of exo-multimodal pretraining data, we propose an Exo-Multimodal Retrieval Enhancer that enables rapid downstream adaptation without updating the TSFM parameters. We further introduce Endo-Modality Proxy Training to train this retrieval module without exogenous multimodal pretraining data. Extensive experiments across unimodal and multimodal benchmarks demonstrate strong forecasting performance in scenarios both with and without exogenous modalities.