arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20554econ.GNq-fin.EC

基于未来数据训练是否有益?预训练模型预测中的前瞻偏差

Does Training on Future Data Pay? Look-Ahead Bias in Forecasting with Pretrained Models

  • Shenzhen University(深圳大学)
  • Xiamen University(厦门大学)
  • University of Science and Technology of China(中国科学技术大学)
  • Chinese Academy of Sciences(中国科学院)

机构由 AI 辅助整理,请以论文原文为准。

Haiqiang Chen, Li Chen, Yunlong Chen, Difang Huang, Bo Zhang

AI总结:

本研究检验预训练金融模型使用未来数据是否虚增预测表现,发现美国训练下事后版本降低准确性,暴露信息集违规而非价值提升。

AI中文摘要:

我们考察了预测起点之后的训练信息是否夸大了金融预测的测量准确性和经济价值。我们评估了五组金融时间序列基础模型,每组模型包含在美国、全球和因子增强训练环境下独立训练的年度版本,覆盖14个股票市场和四个预测期限。滚动比较针对固定预测改变年度版本;固定版本比较在目标窗口跨越其训练截止点时保持版本固定。每个替代预测都与使用相同数值历史和推理协议、对齐预测起点的实时(PIT)基准配对。在美国训练的参考环境中,预测起点之后的版本实质性地修正了信息丰富的PIT预测,但在两种设计中通常降低了准确性。合并的滚动比较在20个美国模型集-期限组合中的18个中产生了更高的均方预测误差。跨越起点的更新平均也劣于同等长度的起点前更新。在使用一个月预测的常见约束分配规则下,美国年化确定性等价收益的中位数暴露减去PIT差异为-1.77个百分点,国际为-2.14个百分点。全球和因子增强训练产生了更多混合的预测效果。精确的平方误差分解表明,当修订的纠错收益超过其均方幅度时,修订提高了准确性;在美国训练下,与PIT误差的对齐通常达不到这一要求。因此,时间暴露确立了信息集违规,而非预测准确性或投资者价值被夸大的充分证据。

英文摘要:

We examine whether post-origin training information inflates the measured accuracy and economic value of financial forecasts. We evaluate five sets of financial time-series foundation models, each comprising independently trained annual vintages under U.S., global, and factor-augmented training environments, across 14 equity markets and four forecast horizons. Rolling comparisons vary the annual vintage for a fixed forecast; fixed-vintage comparisons hold the vintage fixed as target windows move across its training cutoff. Each alternative forecast is paired with an origin-aligned point-in-time (PIT) benchmark using identical numerical histories and inference protocols. In the U.S.-trained reference environment, post-origin vintages materially revise informative PIT forecasts but generally reduce accuracy in both designs. Pooled rolling comparisons yield higher mean squared forecast errors in 18 of 20 U.S. model-set-horizon combinations. The origin-crossing update also performs worse on average than an equally long pre-origin update. Under a common constrained allocation rule using one-month forecasts, median exposed-minus-PIT differences in annualized certainty-equivalent returns are -1.77 percentage points in the United States and -2.14 points internationally. Global and factor-augmented training produce more mixed predictive effects. An exact squared-error decomposition shows that revisions improve accuracy when their error-correcting benefit exceeds their mean squared magnitude; under U.S. training, alignment with PIT errors generally falls short of this requirement. Temporal exposure therefore establishes an information-set violation, not sufficient evidence of inflated predictive accuracy or investor value.

↑