arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12251q-fin.STcs.LG

用于横截面波动率预测的门控残差混合专家模型

Regime-Gated Residual Mixture-of-Experts for Cross-Sectional Volatility Forecasting

  • School of Computing, Montclair State University(蒙特克莱尔州立大学计算机学院)

机构由 AI 辅助整理,请以论文原文为准。

Junyi Ye, Gargi Vijay Borde

中文总结 AI 辅助

本文针对金融波动率的 regime 依赖性问题,提出 RG-ResMoE 架构,将 regime 信息仅用于专家路由,在美、日股票波动率预测中,该模型在精度和稳定性上均优于 MLP。

中文摘要 AI 辅助

金融波动率具有 regime 依赖性,但将 regime 信息纳入神经网络也可能导致训练不稳定。本文探究这类信息应在神经网络横截面波动率预测模型的何处引入。我们采用滚动向前评估框架,对1027只美国股票进行为期5天的已实现波动率预测,该框架在各架构间匹配了信息、模型容量、超参数调优及随机种子。我们提出 RG-ResMoE,一种门控残差混合专家架构,其中 regime 信息仅用于专家路由而非直接预测:基础预测器从股票特征中预测波动率,而门控网络利用 regime 状态变量路由残差修正项。RG-ResMoE 在主要美国研究中,在预测精度和训练稳定性上均持续优于容量匹配的 MLP;在独立日本股票面板上也观察到类似增益。信息整合路径具有决定性:将相同 regime 变量直接附加到预测输入会降低预测性能和训练稳定性,而将其限制在路由门中则可提升精度和风险价值校准效果;硬路由始终逊于软路由。结果表明,在紧凑的神经波动率预测模型中,混合专家模型的主要价值不在于提升模型容量,而在于控制非平稳 regime 信息对预测的影响方式。

英文摘要

Financial volatility is regime dependent, yet incorporating regime information into neural networks can also destabilize training. This paper asks where such information should enter a neural cross-sectional volatility forecasting model. We study five-day realized-volatility forecasts for 1,027 U.S. equities using a rolling walk-forward evaluation framework in which information, model capacity, hyperparameter tuning, and random seeds are matched across architectures. We propose RG-ResMoE, a regime-gated residual mixture-of-experts architecture in which regime information is used only for expert routing rather than for direct forecasting. The base predictor models volatility from stock features, while a gating network uses regime state variables to route residual corrections. RG-ResMoE consistently outperforms a capacity-matched MLP in both forecasting accuracy and training stability in the main U.S. study. Similar gains are observed on an independent Japanese panel. The integration pathway is decisive: appending the same regime variables directly to the forecasting input degrades both predictive performance and training stability, whereas restricting them to the routing gate improves accuracy and Value-at-Risk calibration. Hard routing consistently underperforms soft routing. The results suggest that, in compact neural volatility forecasting models, the primary value of mixture-of-experts models lies less in increasing model capacity than in controlling how nonstationary regime information influences prediction.

↑