arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HAN-Mamba:用于多尺度金融波动率预测的分层选择性状态空间网络

HAN-Mamba: Hierarchical Selective State Space Networks for Multi-Scale Financial Volatility Forecasting

Mihai Bogdan Deaconu, Ioan Daniel Pop

arXiv 2610.10323首次发表:更新:

发表机构

Babeş-Bolyai University(巴比什-博雅伊大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出HAN-Mamba,用Mamba编码器替换HAN-T中的Transformer编码器,在Optiver基准上以更少参数降低RMSPE,并支持更长高频上下文和常数时间推理。

AI 中文摘要

短期已实现波动率预测需要整合以不兼容时间分辨率演化的市场信息,从秒级订单簿动态到周级制度漂移。我们的会议工作引入了HAN-T,一种分层架构,其中特定尺度的Transformer编码器处理短、中、长周期数据流,而学习到的注意力融合器权衡其贡献。本文用选择性状态空间(Mamba)编码器替代二次注意力编码器,仅在融合器中保留注意力,其中输入是三个标记的集合而非长序列。由此产生的混合模型HAN-Mamba通过循环状态总结每个数据流,其依赖于输入的门控匹配波动率的两个结构性质:持久但衰减的记忆和突然的制度转变。在Optiver已实现波动率预测基准上,采用时间感知的五折交叉验证,HAN-Mamba相比HAN-T将平均RMSPE从0.1965降至0.1942,同时参数减少33%。其线性时间编码器还允许将高频上下文从60个桶扩展到240个桶,将误差降至0.1927,而注意力变体在此饱和,并支持推理时的常数时间流式更新。消融实验将增益归因于编码器替换,确认分层先验可跨序列模型族迁移,并表明置换不变的注意力融合器仍是跨尺度集成的正确机制。

英文摘要

Short-horizon realized volatility forecasting requires the integration of market information that evolves at incompatible temporal resolutions, from second-level order book dynamics to weekly regime drift. Our conference work introduced HAN-T, a hierarchical architecture in which scale-specific Transformer encoders process short, mid, and long-horizon streams and a learned attention fuser weighs their contributions. This article replaces the quadratic attention encoders with selective state space (Mamba) encoders while retaining attention only in the fuser, where the input is a three-token set rather than a long sequence. The resulting hybrid, HAN-Mamba, summarizes each stream through a recurrent state whose input-dependent gating matches two structural properties of volatility: persistent but decaying memory and abrupt regime shifts. On the Optiver Realized Volatility Prediction benchmark under time-aware five-fold cross-validation, HAN-Mamba improves mean RMSPE over HAN-T (0.1942 vs. 0.1965) with 33% fewer parameters. Its linear-time encoders further allow the high-frequency context to be extended from 60 to 240 buckets, reducing error to 0.1927 where the attention variant saturates, and support constant-time streaming updates at inference. Ablations attribute the gains to the encoder swap, confirm that the hierarchical prior transfers across sequence-model families, and show that the permutation-invariant attention fuser remains the correct mechanism for cross-scale integration.

Comments16 pages. Accepted for publication in Springer Lecture Notes in Artificial Intelligence (ICAART 2026 Revised Selected Papers). Extended version of the ICAART 2026 paper (DOI: 10.5220/0014264900004052)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑