arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

时间序列基础模型中的预测崩溃现象

Forecast Collapse in Time-Series Foundation Models

Shu Wan, Miles Ma, Hank Zhu, Guangqi Liu, Stephen Wang, Qingsong Wen, Huan Liu

arXiv 2608.14106首次发表:更新:

发表机构

Abel AI Lab; Arizona State University; University of Oxford(阿贝尔人工智能实验室; 亚利桑那州立大学; 牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对1000只美股每小时收益预测出现的预测崩溃现象,本文分析其成因,提出CalibRank目标函数平衡校准与排名,在Finance1K上提升了截面相关性。

AI 中文摘要

在对1000只美国股票的每小时收益进行预测时,我们观察到一种意外现象:预测结果几乎呈平坦状态,且通过截面相关性衡量的股票排名表现很差,我们将此称为预测崩溃。令人惊讶的是,在相同设置下预测交易量时,该现象基本消失。我们在时间序列基础模型(TSFMs)、12种深度学习预测模型及97种公开基准配置中研究了预测崩溃,发现其与目标可预测性密切相关。我们确定了其背后的两个不同原因:低可预测性限制了校准点预测的幅度,而逐系列目标未识别跨系列结构。这些发现揭示了校准-排名权衡:优化平方误差会导致预测结果平坦,而直接优化截面相关性可改善排名,但会使预测幅度增大一个数量级以上。为解决这一权衡,我们引入了CalibRank,一种平衡校准与排名的简单目标函数。在Finance1K数据集上,CalibRank使截面相关性几乎增至原来的三倍,同时保持幅度接近目标,且在所有测试模型上均提高了相关性。我们的结果揭示了传统时间序列评估中的一个盲区:逐系列指标可能隐藏下游决策所需的跨系列结构方面的失败。

英文摘要

When forecasting hourly returns for 1,000 US equities, we observe an unexpected phenomenon: predictions become nearly flat and show poor stock ranking, as measured by cross-sectional correlation. We call this forecast collapse. Surprisingly, the phenomenon largely disappears when forecasting trading volume under the same setting. We investigate forecast collapse across time-series foundation models (TSFMs), twelve deep-learning forecasting models, and 97 public benchmark configurations, and find that it is closely tied to target predictability. We identify two distinct reasons behind it: low predictability limits the amplitude of calibrated point forecasts, while per-series objectives leave cross-series structure unidentified. These findings reveal a calibration-ranking tradeoff: optimizing squared error leads to flat predictions, whereas directly optimizing cross-sectional correlation improves ranking but can inflate forecast amplitude by more than an order of magnitude. To address this tradeoff, we introduce CalibRank, a simple objective that balances calibration and ranking. On Finance1K, CalibRank nearly triples cross-sectional correlation while keeping amplitude close to the target, and improves correlation on all tested models. Our results reveal a blind spot in conventional time-series evaluation: per-series metrics can hide failures in cross-series structure needed by downstream decisions.

Comments27 pages, 3 figures, 5 tables. Dataset: https://huggingface.co/datasets/abel-lab/finance1k

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑