发表机构
GEMS Modern Academy; Vizuara AI Labs(GEMS现代学院; Vizuara人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究以宏观经济预测为压力测试领域,评估五个模型家族在23个国家的表现。结果显示无模型能始终有强大预测性能,约束少的模型如ARIMA和NODE优于约束多的。指出结构先验不匹配数据生成过程时会起反作用,从业者应先测试结构是否有益。
AI 中文摘要
科学机器学习(SciML)方法,如神经常微分方程(NODEs)、物理信息神经网络(PINNs)和通用微分方程(UDEs),当结构先验反映可靠的控制动力学时最为有效。我们研究当这一假设被违反时会发生什么。以宏观经济预测为压力测试领域,我们使用稀疏年度数据、多个时间分割和五个随机种子,对23个国家的五个模型家族ARIMA、LSTM、NODE、PINN和UDE进行评估。结果表明,没有一个评估模型能始终取得强大的预测性能,凸显了低频宏观经济预测的困难。然而,一个明显的相对层次出现了:约束较少的模型,特别是ARIMA和NODE,始终优于约束较多的启发式先验模型如PINN和UDE。我们将此解释为一个诊断结果:当结构先验与数据生成过程不匹配时,它们可能起到错误正则化的作用。我们识别出失败模式,包括先验不匹配、 regime转移、结构断裂和优化不稳定性,并认为SciML从业者在假设更多结构有益之前应测试结构是否有帮助。
英文摘要
Scientific Machine Learning (SciML) methods such as Neural Ordinary Differential Equations (NODEs), Physics-Informed Neural Networks (PINNs), and Universal Differential Equations (UDEs) are most effective when structural priors reflect reliable governing dynamics. We ask what happens when this assumption is violated. Using macroeconomic forecasting as a stress-test domain, we evaluate five model families, ARIMA, LSTM, NODE, PINN, and UDE, across 23 countries using sparse annual data, multiple temporal splits, and five random seeds. Our results show that none of the evaluated models achieve consistently strong forecasting performance, highlighting the difficulty of low-frequency macroeconomic prediction. However, a clear relative hierarchy emerges: less-constrained models, particularly ARIMA and NODE, consistently outperform more-constrained heuristic-prior models such as PINN and UDE. Rather than treating this as a rejection of SciML, we interpret it as a diagnostic result: structural priors can act as misregularizers when they do not match the data-generating process. We identify failure modes including prior misalignment, regime shifts, structural breaks, and optimization instability, and argue that SciML practitioners should test whether structure helps before assuming that more structure is beneficial.