用于掩码扩散训练插补的连续时间序列离散化
Discretizing Continuous Time Series for Imputation with Masked Diffusion Training
浏览论文内容
中文总结 AI 辅助
针对时间序列插补的现有方法存在的局限性,提出MDTIM模型,引入随机离散化技术,在多种基准上的实验显示其鲁棒性和可扩展性优于现有基线方法。
中文摘要 AI 辅助
时间序列插补是可靠时间序列分析的关键领域,但由于真实数据的复杂时间动态和噪声,它仍然具有挑战性。然而,现有方法存在两个局限性:缺失值和观测值被嵌入同一表示空间,未进行显式结构分离;基于连续扩散的方法被训练用于预测添加的噪声而非原始信号。为解决这些问题,我们提出了掩码扩散时间序列插补模型(Masked Diffusion Time-series Imputation Model,MDTIM),该模型利用掩码扩散模型的训练范式完成插补任务。MASK 标记与有效观测值结构正交,模型直接预测原始值,使表示和学习目标自然与插补任务对齐。为弥合离散掩码扩散与时间序列的连续序数性质之间的差距,我们进一步引入了随机离散化(Stochastic Discretization),该方法将连续值映射到感知序数的标记,同时保留连续动态。我们在多种基准上开展的实验证实,MDTIM 具备出色的鲁棒性和可扩展性,在各类缺失场景下均持续优于最先进的确定性和生成式基线方法。
英文摘要
Time series imputation is a crucial area for reliable time series analysis, yet it remains challenging due to the complex temporal dynamics and noise of real-world data. Existing approaches, however, exhibit two limitations: missing and observed values are embedded within the same representation space without explicit structural separation, and continuous diffusion-based methods are trained to predict added noise rather than the original signal. To address these, we propose the Masked Diffusion Time-series Imputation Model (MDTIM), which leverages the training paradigm of masked diffusion model for imputation tasks. The MASK token is structurally orthogonal to valid observations, and the model directly predicts the original values, naturally aligning both the representation and the learning objective with the imputation task. To bridge the gap between discrete masked diffusion and the continuous, ordinal nature of time series, we further introduce Stochastic Discretization, which maps continuous values to ordinal-aware tokens while preserving continuous dynamics. Our experiments on diverse benchmarks confirm that MDTIM achieves superior robustness and scalability, consistently outperforming state-of-the-art deterministic and generative baselines across various missing scenarios.
发表机构
- Seoul National University(首尔大学)
机构由 AI 辅助整理,请以论文原文为准。