arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过多源异常融合在标签稀缺情况下实现极端波动预警

Extreme Volatility Warning under Label Scarcity via Multi-Source Anomaly Fusion

Jin Qian, Zhangzhi Xiong, Mingrui Li, Zhen Liu

arXiv 2607.23682首次发表:更新:

发表机构

ShanghaiTech University(上海科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对金融市场极端波动预警中标签稀缺问题,提出半监督框架AAMSF,结合多源数据的异常分数与轻量级岭分数融合,还引入时间扩展T - AAMSF,在沪深300指数上取得良好效果,揭示源不对称性,给出标签稀缺下金融风险预警的设计原则。

AI 中文摘要

极端市场波动的早期预警是金融风险管理的核心,但可操作事件罕见、非平稳且常由外部信息冲击触发。在沪深300指数设置中,791个训练日里仅有约80个正样本,导致强监督多源模型不稳定。首先分析了一个100K参数的分层文本信号融合模型(HTSF),发现增加参数化在低标签情况下有害。基于此失败,提出了AAMSF(异常增强多信号融合),这是一个半监督框架,将市场指标、GDELT事件、中文财经新闻和英文媒体上的孤立森林异常分数与轻量级岭分数融合相结合。还引入了T - AAMSF,用于多日异常积累的时间扩展。在沪深300指数(2018 - 2023)上,AAMSF的测试AUC - ROC达到0.680,优于最强无监督基线(0.630)和神经基线(0.588),而T - AAMSF将PR - AUC提高到0.291。消融显示出强烈的源不对称性:GDELT和国内财经新闻提供互补风险信号,而英文媒体持续降低性能,且在验证噪声下学习的权重不可靠。这些结果表明了标签稀缺的金融风险预警的经验设计原则:强大的异常几何结构和源可靠性可能比监督表示能力更重要。

英文摘要

Early warning of extreme market volatility is central to financial risk management, but actionable events are rare, nonstationary, and often triggered by exogenous information shocks. In our CSI~300 setting, only $\sim$80 positive samples are observed across 791 training days, making heavily supervised multi-source models unstable. We first analyze a 100K-parameter hierarchical text-signal fusion model (HTSF) and find that added parameterization hurts in this low-label regime. Motivated by this failure, we propose \textbf{AAMSF} (Anomaly-Augmented Multi-Signal Fusion), a semisupervised framework that combines Isolation Forest anomaly scores over market indicators, GDELT events, Chinese financial news, and English media with lightweight Ridge score fusion. We further introduce \textbf{T-AAMSF}, a temporal extension for multi-day anomaly accumulation. On CSI~300 (2018--2023), AAMSF achieves test AUC-ROC \textbf{0.680}, outperforming the strongest unsupervised baseline (0.630) and neural baseline (0.588), while T-AAMSF improves PR-AUC to 0.291. Ablations reveal strong source asymmetry: GDELT and domestic financial news provide complementary risk signals, whereas English media consistently reduces performance, and learned weighting is unreliable under validation noise. These results suggest an empirical design principle for label-scarce financial risk warning: robust anomaly geometry and source reliability can matter more than supervised representation capacity.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑