发表机构
School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences; State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences; School of Intelligent Systems Engineering, Sun Yat-sen University; School of Computer Science and Technology, Huazhong University of Science and Technology; School of Artificial Intelligence, University of Chinese Academy of Sciences; College of Computer Science, Chongqing University(中国科学院大学先进交叉科学学院; 中国科学院自动化研究所多模态人工智能系统国家重点实验室; 中山大学智能系统工程学院; 华中科技大学计算机科学与技术学院; 中国科学院大学人工智能学院; 重庆大学计算机学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对媒体桥接时间序列预测的范式壁垒,提出MIDAPN统一主干网络,经多数据集对比验证其具备持续优越性与广泛兼容性。
AI 中文摘要
媒体桥接时间序列预测正扩展至涵盖传统“多变量”及新兴“多模态”(例如通过文本辅助)场景。现有时间序列预测(TSF)模型仍依赖特定范式的关系、融合与时序模块,阻碍了适用于数值数据与预对齐叙事流场景的通用预测主干网络。为探索此问题,我们提出多媒体身份感知棱镜网络(MIDAPN),一种基于媒体通用图适配与自动时序学习的统一时空预测主干网络:(1)经媒体预对齐后,我们的多媒体身份感知图(MIDAG)通过静态本质、动态行为与潜在共性重新审视身份,生成跨媒体扩展变量特定依赖的亲和关系;上下文身份调制(CIM)进一步优化判别性聚合。(2)我们开发谱棱镜卷积(SPConv)以自动执行分层时序分析,平衡粗粒度趋势与细粒度细节,同时其自适应搜索引导为时序维度重构配置了规模高效的架构。这些解耦却协同的组件共同解决媒体身份解缠与时序尺度不匹配问题。涉及16个SOTA TSF模型、13个“多变量”数据集及12个“多模态”数据集的综合评估,以及针对14个时序基础模型与融合预训练语言模型的定向长上下文对比,均证明MIDAPN具备持续优越性与广泛的共享主干兼容性。代码可在指定链接获取。
英文摘要
Media-bridged time series forecasting is expanding to encompass traditional "multivariate" and emerging "multimodal" (e.g., through textual assistance). Existing Time Series Forecasting (TSF) models still rely on paradigm-specific relation, fusion, and temporal modules, hindering a common forecasting backbone across numerical and pre-aligned narrative-flow settings. To explore this, we propose the Multimedia Identity-Aware Prism Network (MIDAPN), a unified spatiotemporal forecasting backbone based on media-general graph adaptation and automatic temporal learning: (1) Following media pre-alignment, our Multimedia Identity-Aware Graph (MIDAG) revisits identity through static essence, dynamic behavior, and latent commonality, inducing affinities that extend variable-specific dependencies across media. Contextual Identity Modulation (CIM) further refines discriminative aggregation. (2) We develop Spectral Prism Convolution (SPConv) to automatically perform hierarchical temporal analysis, balancing coarse trends and fine-grained details. Meanwhile, its Adaptive Search Guidance configures a scale-efficient architecture for temporal-dimension reconstruction. These decoupled yet synergistic components jointly address media identity disentanglement and temporal-scale mismatch. Comprehensive evaluations involving 16 SOTA TSF models across 13 "multivariate" and 12 "multimodal" datasets, alongside targeted long-context comparisons against 14 time series foundation models and fused pretrained language models, demonstrate MIDAPN's consistent superiority and broad shared backbone compatibility. The code is available at \href{https://github.com/leijieruilq/MIDAPN/tree/main}{https://github.com/MIDAPN}.