基于通道依赖状态空间模型的多变量时间序列预测
Channel-Dependent State Space Model for Multivariate Time Series Forecasting
浏览论文内容
中文总结 AI 辅助
提出Chameleon,一种通道依赖状态空间模型,通过结合卡尔曼滤波实现跨变量交互,在多个基准上取得最优预测精度。
中文摘要 AI 辅助
多变量时间序列预测(MTSF)在许多现实领域中都至关重要。现有的深度学习方法分为两种范式,各有明显局限:通道独立(CI)方法无条件忽略跨变量依赖,仅对时间动态建模;而通道依赖(CD)方法虽同时考虑两者,但通常依赖架构上的折衷来缓解过拟合和计算开销。为此,我们提出了Chameleon,一种专门的CD状态空间模型(SSM),它能够在变量间实现数据依赖的细粒度交互,同时其计算复杂度随变量数量线性增长。通过将选择性SSM与卡尔曼滤波器连接,我们利用前者缺失的测量更新来进行跨变量建模,同时保留SSM主干以实现稳健的时间建模。我们进一步识别了GatedDeltaNet对时间序列的有利归纳偏置,将其适配为我们的主干,并通过额外技术改进泛化能力,包括一种先前未探索的可逆实例归一化的随机扰动。在强依赖的ODE和PEMS数据集上,Chameleon在所有设置中均取得了最佳MSE和MAE,而其CI消融和先前的CD方法平均MSE高出61-178%。在28个标准基准设置中,Chameleon在至少27个和22个案例中分别取得了优于每个基线的MSE和MAE。在Traffic和ETT上的训练时间和峰值内存分析进一步证明了其在不同变量数量下的竞争性效率和良好的内存可扩展性。
英文摘要
Multivariate time series forecasting (MTSF) is critical across many real-world domains. Existing deep learning approaches fall into two paradigms with distinct limitations: channel-independent (CI) methods unconditionally ignore cross-variable dependencies and model only temporal dynamics, while channel-dependent (CD) methods consider both but typically rely on architectural compromises to mitigate overfitting and computational overhead. We therefore propose Chameleon, a specialized CD state space model (SSM) that enables data-dependent, fine-grained interactions across variables while scaling linearly with their number. By connecting selective SSMs with the Kalman filter, we leverage the missing measurement update in the former for cross-variable modeling while preserving the SSM backbone for robust temporal modeling. We further identify favorable inductive biases of GatedDeltaNet for time series, adapt it as our backbone, and improve generalization through additional techniques, including a previously unexplored stochastic perturbation of reversible instance normalization. On strongly dependent ODE and PEMS datasets, Chameleon achieves the best MSE and MAE across all settings, while its CI ablation and prior CD methods incur 61-178% higher MSE on average. Across 28 standard benchmark settings, Chameleon also achieves better MSE and MAE than each baseline in at least 27 and 22 cases, respectively. Training-time and peak-memory analyses on Traffic and ETT further demonstrate competitive efficiency and favorable memory scalability across different variable counts.
发表机构
- Massachusetts Institute of Technology(麻省理工学院)
- National Taiwan University(国立台湾大学)
机构由 AI 辅助整理,请以论文原文为准。