arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LiveHouse-TS:面向时间序列基础模型的开放世界动态基准

LiveHouse-TS: An Open-world Living Benchmark for Time Series Foundation Models

Haomin Wen, Ziyu Zhou, Qingxiang Liu, Siru Zhong, Yuxuan Liang

arXiv 2608.17299首次发表:更新:

发表机构

The Hong Kong University of Science and Technology (Guangzhou); Shanghai Innovation Institute(香港科技大学(广州); 上海创新研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出首个面向时间序列基础模型(TSFMs)的开放世界动态基准LiveHouse-TS,通过序贯评估发现静态基准的模型排名在动态环境下会剧烈变化,为评估TSFMs的长期鲁棒性提供了新框架。

AI 中文摘要

时间序列基础模型(TSFMs)是近期兴起的、极具潜力的跨域零样本预测范式。然而,现有评估协议主要依赖具有固定历史测试窗口的静态基准,这些基准虽能提供有价值的基线快照,却仅评估模型在固定历史数据上的平均性能,无法捕捉模型在持续演化的真实世界环境中的表现——这类环境具有季节变化、分布偏移和意外事件等特征。为填补这一空白,我们推出LiveHouse-TS,这是首个面向TSFMs的开放世界动态基准基础设施。LiveHouse-TS通过在开放世界环境中对真实未来数据进行序贯评估,将时间序列基准从快照精度转向连续时间有效性。该基础设施并非一次性的排行榜,而是旨在探索重要长期科学问题的连续时间序列基础设施:模型排名能否在长期内保持稳定?哪些模型在分布偏移下仍保持真正的鲁棒性?对11个领域的17个数据集开展的广泛流处理评估表明,在动态协议下,静态排名会发生剧烈的重新洗牌。

英文摘要

Time Series Foundation Models (TSFMs) have recently emerged as a highly promising paradigm for cross-domain zero-shot forecasting. However, existing evaluation protocols predominantly rely on static benchmarks with fixed historical test windows. While these benchmarks provide a valuable baseline snapshot, they evaluate an average performance on a fixed history, failing to capture how models behave in continuously evolving real-world environments characterized by seasonal variations, distribution shifts, and unexpected events. To bridge this gap, we introduce LiveHouse-TS, the first open-world living benchmark infrastructure for TSFMs. By evaluating models prequentially on real future data in open-world environments, LiveHouse-TS shifts time series benchmarking from snapshot accuracy to continuous temporal validity. Rather than acting as a one-off leaderboard, our infrastructure serves as a continuous time series infrastructure designed to explore vital, long-term scientific questions: Can model rankings be maintained over the long term? Which models remain genuinely robust under distribution shifts? Extensive streaming evaluations across 11 domains with 17 datasets demonstrate that static rankings undergo a dramatic reshuffling under a live protocol.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑