AI 中文总结
针对不规则时间序列预测缺乏统一基准的问题,提出BITS基准,涵盖11个数据集和多种指标,发现方法性能随不规则性变化且指标选择影响排名,强调多维评估。
AI 中文摘要
尽管不规则时间序列预测近期取得了进展,该领域仍缺乏一个用于公平和全面评估的统一基准。现有评估通常在一组有限的数据集上进行,实验协议不一致,且主要采用基于误差的指标,这使得在不同设置下难以公平且全面地比较和评估方法。为消除这些限制并加速进展,我们提出了BITS,一个标准化、可复现且可扩展的基准,用于推进不规则时间序列预测的研究。BITS涵盖了来自九个领域的十一个数据集,具有多样的不规则性特征,并根据缺失率、缺失模式复杂度、采样不规则性和偏度对数据集进行刻画。此外,它提供了一个统一的数据预处理、模型集成、评估和报告流程。它在一致设置下容纳规则和不规则时间序列预测方法,包括时间序列基础模型,并整合了基于误差和非基于误差的评估指标。研究发现,方法性能在不同不规则性特征间差异显著,没有单一建模策略始终占优。我们还发现,使用基于误差或非基于误差的指标可能产生不同的模型排名,凸显了多维评估的必要性。代码可在该https URL找到。
英文摘要
Despite recent progress in irregular time series forecasting, the field still lacks a unified benchmark for fair and comprehensive evaluation. Existing evaluations are often conducted on a limited set of datasets with inconsistent experimental protocols and predominantly error-based metrics, rendering it difficult to compare and assess methods fairly and comprehensively across diverse settings. To eliminate these limitations and accelerate progress, we propose BITS, a standardized, reproducible, and extensible benchmark for advancing research on irregular time series forecasting. BITS covers eleven datasets from nine domains with diverse irregularity characteristics, and it characterizes the datasets according to their missing rate, missing pattern complexity, sampling irregularity, and skewness. Further, it offers a unified pipeline for data preprocessing, model integration and evaluation, and reporting. It accommodates regular and irregular time series forecasting methods, including time series foundation models, under consistent settings, incorporating both error-based and non-error-based evaluation metrics. Findings include that method performance varies substantially across irregularity characteristics, with no single modeling strategy consistently dominating. We also find that using error-based or non-error-based metrics can yield different model rankings, highlighting the need for multi-dimensional evaluation. The code can be found at https://anonymous.4open.science/r/BITS-8F2E/.