发表机构
Univ. Grenoble Alpes; CNRS; Grenoble INP; LIG; Savoye(格勒诺布尔阿尔卑斯大学; 法国国家科学研究中心; 格勒诺布尔理工学院; 信息与信号处理实验室; 萨沃伊公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出自适应可逆差分(AdaRDiff)即插即用模块,通过可学习权重的加权差分稳定残差,实现长程时间序列预测精度提升,可改进多种骨干网络且加速明显。
AI 中文摘要
可靠的长程时间序列预测是重要但困难的问题,趋势和季节引入的复杂时间结构对基于学习的预测模型构成挑战。差分(通过减去邻近过去值以消除此类结构)是经典解决方案,但它依赖手动选择的阶数和周期,因此在近期深度架构中基本未被采用。我们提出自适应可逆差分(AdaRDiff),这是一种广义差分方法,使用可学习权重通过与先前时间步的加权差分来简化序列。这会产生稳定的残差,在此残差上执行预测,之后通过自回归方式恢复已移除的分量以重构预测,通过单个算子联合捕获趋势和季节。该重构可导出闭式卷积表达式,能在GPU上并行化,相比朴素循环可实现高达33.7倍的加速。此外,我们采用两阶段训练方案,将时间结构发现与重构学习分离,这是使用线性预测模型时梯度理论分析所建议的。AdaRDiff在涵盖电力、天气、交通和能源的8个基准测试中达到了最先进的预测精度,且参数成本可忽略不计。此外,它被设计为即插即用模块:集成AdaRDiff可在绝大多数情况下改进8种不同的骨干网络,从线性模型到Transformer,使用线性骨干网络时提升高达25.9%,使用iTransformer时提升高达18.3%。
英文摘要
Reliable long-horizon time series forecasting is an important yet difficult problem. Trends and seasonality introduce complex temporal structure that challenges learning-based forecasting models. Differencing, which subtracts nearby past values to remove such structure, is the classical remedy, but its reliance on hand-picked orders and periods has kept it largely absent from recent deep architectures. We propose \textbf{\underline{Ada}}ptive \textbf{\underline{R}}eversible \textbf{\underline{Diff}}erencing \textbf{(AdaRDiff)}, a generalized differencing approach that uses learnable weights to simplify the series through weighted differencing with previous time instants. This yields stabilized residuals on which forecasting is performed, after which the removed components are restored autoregressively to reconstruct the forecast, capturing trend and seasonality jointly through a single operator. This reconstruction admits a closed-form convolutional expression, which parallelizes on GPU and yields up to $33.7\times$ speedup over the naive recurrence. We furthermore rely on a two-phase training schedule that separates temporal structure discovery from reconstruction learning, as suggested by a theoretical analysis of the gradient when using a linear forecasting model. AdaRDiff attains state-of-the-art forecast accuracy across eight benchmarks spanning electricity, weather, traffic, and energy, at negligible parameter cost. Furthermore, it is designed as a plug-and-play module: integrating AdaRDiff improves eight diverse backbones, from linear models to Transformers, in the large majority of cases, by up to $25.9\%$ with a linear backbone and $18.3\%$ with iTransformer.