FinVerse:金融时间序列基准
FinVerse: Financial Time-Series Benchmark
浏览论文内容
中文总结 AI 辅助
本研究推出金融时间序列预测基准FinVerse,其针对43个公开时间序列基础模型的分析显示,通用预测标准下的优异表现未必对应实用金融预测,凸显了领域感知基准的必要性。
中文摘要 AI 辅助
随着时间序列基础模型的兴起,对能够以有意义的方式评估其预测能力的基准的需求日益重要。现有的时间序列预测基准提供了有用的标准化比较,但它们通常使用统一的基于误差的指标评估异构序列。在这类指标下表现出色并不一定意味着模型的预测能支持各领域最佳的现实决策。例如,在股票预测中,与单纯最小化逐点预测误差相比,正确预测价格涨跌与实现的收益更直接相关。为此,我们推出FinVerse,这是一个金融领域的时间序列预测基准,朝着更现实的评估迈出了第一步。发布的FinVerse数据制品包含116897个金融时间序列,共1.711亿个观测值,其中基于与金融决策的经济相关性,选择了60232个序列(共1740万个观测值)作为评估目标。与主要强调统一点预测或概率准确性的通用预测基准不同,FinVerse定义了包含78个评估指标的11个指标家族,并根据每个时间序列的潜在经济含义为其分配最合适的评估指标。我们对43个公开的时间序列预测基础模型的分析表明,在通用预测标准下的出色表现不一定能转化为有用的金融预测。这一发现凸显了对领域感知基准的需求,这类基准需在更贴近现实决策的目标下评估模型。
英文摘要
As time-series foundation models have emerged, the need for benchmarks that can evaluate their forecasting ability in meaningful ways has become increasingly important. Existing time-series forecasting benchmarks provide useful standardized comparisons, but they often evaluate heterogeneous series with uniform error-based metrics. Strong performance under such metrics does not necessarily imply that a model's forecasts will support the best real-world decisions across domains. For example, in stock forecasting, correctly predicting whether a price will rise or fall can be more directly relevant to realized returns than minimizing point-wise forecast error alone. To this end, we introduce FinVerse, a finance-domain time-series forecasting benchmark that takes a first step toward more realistic evaluation. The released FinVerse data artifact contains 116,897 financial time series with 171.1M observations, of which 60,232 series with 17.4M observations are selected as evaluated targets based on their economic relevance to financial decisions. Unlike generic forecasting benchmarks that primarily emphasize uniform point-forecast or probabilistic accuracy, FinVerse defines 11 metric families comprising 78 evaluation metrics and assigns the most appropriate evaluation metrics to each individual time series based on its underlying economic meaning. Our analysis of 43 public time-series forecasting foundation models shows that strong performance under generic forecasting criteria does not necessarily translate into useful financial forecasts. This finding highlights the need for domain-aware benchmarks that evaluate models under objectives closer to real-world decision making.