arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

雾中的预测:机器学习相对于菲利普斯曲线的实时数据与修正数据证据

Forecasting in the Fog: Real-Time versus Revised-Data Evidence on Machine Learning's Edge over the Phillips Curve

Louis Agyekum, Obed Obese

arXiv 2608.09033首次发表:更新:

AI 中文总结

本文基于美国2000-2026年数据,对比实时与修正数据下ML模型和传统模型的通胀预测表现,发现ML优势在实时数据下减弱,SHAP特征重要性排名存在事后偏差,RSI可用于审计ML通胀预测。

AI 中文摘要

机器学习(ML)通胀预测几乎普遍基于完全修正后的数据训练,然而实时预测者从未能获得此类数据,且报告的特征重要性通常是通过样本内计算得出的,这混淆了预测相关性与回溯拟合度。本文探究Agyekum(2026)中所记录的ML相对于菲利普斯曲线的优势,在模型基于实时(ALFRED) vintage而非修正序列进行训练和评估时是否仍然存在,以及SHAP特征重要性排名是否为样本内估计的产物。利用2000-2026年美国的失业率、CPI、PCE通胀、就业人数、实际GDP以及10年期-2年期国债利差数据,构建了4种传统模型(随机游走、AR(1)、菲利普斯曲线、ADL-OLS)和4种ML模型(随机森林、梯度提升、弹性网、SVR)的vintage一致面板,在3、6、12个月的预测期内递归重新估计(分别产生208、206、204个预测值)。实时/修正数据的精度差异很小,除了6个月期的一个例外(梯度提升vs菲利普斯曲线,DM=-1.671,p=0.097),在Diebold-Mariano检验下无显著差异;仅梯度提升在较长预测期表现出持续的正预测能力。随机游走仍是强劲的短期基准,与Agyekum等人(2026)关于汇率的谜题一致。采用滚动向前的样本外SHAP,基于修正数据的随机森林将主要重要性赋予PCE通胀(平均|SHAP|=0.778,9个特征中排名第1),而基于实时数据时,其赋予PCE的重要性可忽略不计(0.039,排名第6),转而依赖当期CPI(0.834,排名第1,PCE排名第2为0.223)。这一20倍的波动幅度大于样本内估计值,无法被点预测指标察觉,表明模型对PCE的依赖在很大程度上是一种事后产物。RSI按模型和预测期汇总了精度差距,对ML通胀预测的审计具有启示意义。

英文摘要

ML inflation forecasts are almost universally trained on fully revised data, even though real-time forecasters never have such data, and reported feature importances are typically computed in-sample, conflating predictive relevance with retrospective fit. This paper asks whether the ML advantage over the Phillips curve documented in Agyekum (2026) survives when models are trained and evaluated on real-time (ALFRED) vintages rather than revised series, and whether SHAP feature-importance rankings are an artifact of in-sample estimation. Using 2000-2026 U.S. data on unemployment, CPI and PCE inflation, payrolls, real GDP, and the 10-year-2-year Treasury spread, vintage-consistent panels are built for four traditional models (random walk, AR(1), Phillips curve, ADL-OLS) and four ML models (Random Forest, Gradient Boosting, Elastic Net, SVR), re-estimated recursively at 3-, 6-, and 12-month horizons (208, 206, 204 forecasts). Real-time/revised accuracy differences are small and, apart from one exception at 6 months (Gradient Boosting vs. Phillips curve, DM = -1.671, p = 0.097), indistinguishable under Diebold-Mariano tests; Gradient Boosting alone shows consistent positive skill at longer horizons. The random walk remains a strong short-horizon benchmark, consistent with the puzzle in Agyekum et al. (2026) for exchange rates. Using walk-forward, out-of-sample SHAP, a Random Forest on revised data assigns dominant importance to PCE inflation (mean |SHAP| = 0.778, rank 1 of 9), while on real-time data it assigns PCE negligible importance (0.039, rank 6), relying instead on current CPI (0.834 vs. 0.223). This twenty-fold swing, larger than the in-sample estimate, is invisible to point-forecast metrics and shows the model's PCE reliance is substantially a hindsight artifact. An RSI summarizes the accuracy gap by model and horizon, with implications for auditing ML inflation forecasts.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑