发表机构
LIVIA, Department of Software and IT Engineering; École de technologie supérieure (ÉTS); Center for Informatics (CIn); Universidade Federal de Pernambuco (UFPE)(LIVIA,软件与IT工程系; 高等技术学院; 信息学中心; 伯南布哥联邦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究揭示现代预测基准因忽视预处理(如差分)而产生结构性偏差,通过预处理感知基准测试证明优化预处理可显著提升简单模型性能,使其能与复杂模型竞争。
AI 中文摘要
尽管已有文献强调了预处理在预测准确性中的关键作用,但这一阶段在当前研究中仍被大大忽视。现代基准测试通常采用简单的缩放,未能考虑处理非平稳性所需的关键变换,如差分。这一遗漏造成了显著的结构性预处理偏差,偏向于具有内置数据处理的模型,同时掩盖了更简单架构的真正潜力。我们通过一个预处理感知的基准测试来研究这一效应,该基准在29,000个M4时间序列上,跨16个可逆预处理流程评估了11个预测模型。我们的结果将预处理确定为预测性能的关键驱动因素。针对每个序列优化预处理,在所有评估模型上带来了约27%至87%的增益,其中缺乏内化预处理的架构获得了最显著的改进。这使得更简单的架构在现代预测基准测试中能够与复杂的最先进模型高度竞争。该基准测试的所有资源和实验结果存储在一个全面的元数据集中,以支持未来的元学习任务。
英文摘要
While established literature underscores the pivotal role of preprocessing in forecasting accuracy, this stage remains largely overlooked in current research. Modern benchmarks typically resort to simple scaling, failing to account for critical transformations required to address nonstationarity, such as differencing. This omission creates a significant structural preprocessing bias that favors models with built-in data treatments while obscuring the true potential of simpler architectures. We study this effect through a preprocessing-aware benchmark that evaluates 11 forecasting models across 16 reversible preprocessing pipelines on 29,000 M4 time series. Our results identify preprocessing as a key driver of forecasting performance. Optimizing preprocessing per series yields gains of approximately 27\% to 87\% across all evaluated models, with architectures lacking internalized preprocessing experiencing the most substantial improvements. This allows simpler architectures to become highly competitive with complex, state-of-the-art models in modern forecasting benchmarks. All resources and experimental results from this benchmark are stored in a comprehensive metadataset to support future metalearning tasks.
Comments29 pages, 9 figures. Under Review