发表机构
Southeast University; Nanyang Technological University; Timecho Ltd.(东南大学; 南洋理工大学; 天谋科技)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出混合注意力模型,通过时间感知补丁编码与时间偏置注意力处理不规则多变量时间序列,在三个IMTS基准上取得最优零样本性能。
AI 中文摘要
时间序列基础模型(TSFMs)近期在多种预测任务中展现了令人印象深刻的零样本性能。然而,现实世界的决策常常依赖于不规则多变量时间序列(IMTS),其中观测间隔不一致、变量间异步采样与信息性缺失并存。现有的TSFMs处理此类输入时,要么通过插值引入虚假值,要么采用基于索引的位置编码而忽略连续时间。在遵循原始IMTS模式的基础模型方面仍存在空白。本文提出一种混合注意力模型,学习统一的IMTS预测时间感知补丁表示。我们首先设计了一种时间感知补丁编码,将可变数量的补丁内时间戳映射为固定大小的嵌入,在不借助插值的情况下为不规则补丁生成统一格式。接着引入时间偏置注意力机制,将补丁间时间错位和异步跨通道依赖校准为辅助注意力偏移。最后,在仅解码器Transformer骨干之上,采用混合因果掩码,在保持预测范围严格自回归的同时,对历史上下文保留双向完整视图。为支持不规则设置下的大规模预训练,我们还整理了VersaTSA,一个包含300亿观测值、保留其来源原生采样稀疏性的档案库。在三个IMTS基准和一个标准规则MTS基准上的实验表明,我们的模型在IMTS上达到最先进的零样本性能,并在迁移至规则预测时保持竞争力。
英文摘要
Time series foundation models (TSFMs) have recently delivered impressive zero-shot performance across diverse forecasting tasks. However, real-world decision-making frequently relies on \emph{irregular multivariate time series} (IMTS), where inconsistent inter-observation intervals and asynchronous sampling across variables coexist with informative missingness. Existing TSFMs handle such inputs either through imputation that injects spurious values or through index-based positional encodings that ignore continuous time. There is still a gap in the foundation model that follows the original IMTS patterns. In this paper, we propose a hybrid attention model that learns a unified time-aware patch representation for IMTS forecasting. We first design a \emph{time-aware patch encoding} that maps a variable number of intra-patch timestamps into a fixed-size embedding, producing a uniform format for irregular patches without resorting to imputation. We then introduce a \emph{time bias attention} mechanism that calibrates inter-patch temporal misalignment and asynchronous cross-channel dependencies as auxiliary attention offset. Finally, on top of a decoder-only Transformer backbone, we adopt a \emph{hybrid causal mask} that preserves a bidirectional full view over the historical context while keeping the forecast horizon strictly autoregressive. To support large-scale pretraining under irregular settings, we also curate VersaTSA, an archive of $30$B observations that retains the native sampling sparsity of its sources. Experiments on three IMTS benchmarks and a standard regular-MTS benchmark show that our model achieves state-of-the-art zero-shot performance on IMTS and remains competitive when transferred to regular forecasting.