AI 中文总结
本文提出模型无关的潜在制度偏差审计框架,用于评估波动率预测在不同市场制度下的可靠性,发现聚合准确率高的模型仍存在制度特定偏差和尾部预测不足,为波动率预测评估提供新方向。
AI 中文摘要
波动率预测通常采用RMSE和MAE等聚合准确率指标进行评估,但这些指标可能隐藏对风险管理至关重要的条件性失效。本文提出一种模型无关的审计框架,用于评估波动率预测在潜在市场制度下是否保持可靠。我们学习市场状态窗口的时间序列表示,仅使用训练信息将其聚类为制度,对样本外数据分配制度,并比较聚合预测行为与制度条件偏差、尾部预测不足以及对预测不足敏感的经济损失。将该框架应用于加密货币和ETF资产的日波动率预测,审计显示,具有竞争力聚合准确率的模型仍可能表现出显著的制度特定偏差和严重的尾部预测不足。结果表明,波动率预测不仅应通过平均误差评估,还应评估预测变得不可靠的场景和方式。我们的框架将预测评估从“哪个模型平均最准确”转变为识别看似准确的预测在哪些市场制度下会出现条件性失效。可复现性:此httpsURL
英文摘要
Volatility forecasts are commonly evaluated with aggregate accuracy metrics such as RMSE and MAE, but these metrics can hide conditional failures that matter for risk management. This paper proposes a model-agnostic audit framework for evaluating whether volatility forecasts remain reliable across latent market regimes. We learn time-series representations of market-state windows, cluster them into regimes using only training information, assign regimes out of sample, and compare aggregate forecast behavior with regime-conditional bias, tail-underprediction, and underprediction-sensitive economic losses. Applied to daily volatility forecasting across cryptocurrency and ETF assets, the audit shows that models with competitive aggregate accuracy can still exhibit substantial regime-specific bias and severe tail underprediction. The results suggest that volatility forecasting should be evaluated not only by average error, but also by where and how forecasts become unreliable. Our framework shifts forecast evaluation from asking which model is most accurate on average to identifying the market regimes in which apparently accurate forecasts fail conditionally. Reproducibility: https://github.com/arthurchagas1/Latent-Regime-Bias-Auditing-for-Volatility-Forecasting
CommentsAccepted for publication at the IEEE Conference on Computational Intelligence for Financial Engineering & Economics (CIFEr 2026)