arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越均方误差:重新审视不规则时间序列预测的评估指标与基准测试

Beyond MSE: Rethinking the Evaluation Metric and Benchmarking for Irregular Time Series Forecasting

Rongwen Li, Haixin Xie, Xiao Wang, Changjian Chen

arXiv 2608.17293首次发表:更新:

发表机构

College of Computer Science and Electronic Engineering, Hunan University(湖南大学计算机与通信学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对不规则时间序列预测中MSE评估存在偏差的问题,提出CSE指标并构建多类型数据集基准,验证其能更准确评估模型连续时间预测性能。

AI 中文摘要

现有不规则时间序列预测研究主要聚焦于模型设计,而对评估指标的研究仍不充分;现有基准通常采用均方误差(MSE)作为评估指标。本文表明,在不规则预测中,MSE不仅由模型预测结果决定,还受特定样本的时间戳采样分布影响,导致对模型连续时间预测性能的评估存在偏差。为解决该问题,本文提出连续时间均方误差(CSE),该指标采用重要性加权以消除时间戳采样分布的影响;本文从理论上证明,CSE关于连续时间风险的渐近估计误差不大于MSE的对应误差。最后,本文构建了涵盖合成、半合成及8个真实世界数据集的系统基准,以验证CSE的有效性并系统评估模型的连续时间预测性能。实验显示,CSE相比MSE能更准确地恢复连续时间风险,而仅依赖MSE可能无法充分反映模型在真实场景中的连续时间预测性能;本文代码可通过该https链接获取。

英文摘要

Existing research on irregular time-series forecasting has primarily focused on model design, while evaluation metrics remain insufficiently studied. Existing benchmarks typically use mean squared error (MSE) as the evaluation metric. We show that, in irregular forecasting, MSE is determined not only by the model prediction but also by the sample-specific timestamp sampling distributions, leading to a biased assessment of the models' continuous-time predictive performance. To address this issue, we propose the Continuous-time Squared Error (CSE), which employs importance weighting to eliminate the influence of the timestamp sampling distributions. We further theoretically prove that CSE's asymptotic estimation error with respect to continuous-time risk is no greater than that of MSE. Finally, we construct a systematic benchmark covering synthetic, semi-synthetic, and eight real-world datasets to validate the effectiveness of CSE and systematically evaluate models' continuous-time predictive performance. Experiments show that CSE can recover continuous-time risk more accurately than MSE, while relying solely on MSE may not fully reflect models' continuous-time predictive performance in real-world scenarios. Our code can be obtained at https://github.com/hnu-vis/ITS-Bench.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑