arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向原始日内K线的横截面股票收益率排序的紧凑选择性状态空间模型

A Compact Selective State-Space Model for Cross-Sectional Stock Return Ranking from Raw Intraday Bars

Mingju Chen, Enze Zhang, Annan Li, Yui Lo, Xiaomin Yuan, Kaiming Yu, Jinhui Ren, Yuanhang Liu

arXiv 2608.28060首次发表:更新:

发表机构

Famou Agent Team, Baidu AI Cloud; Qiuzhen College, Tsinghua University; Institute for AI Industry Research (AIR), Tsinghua University, Beijing, China(百度智能云 发悟智能体团队; 清华大学 邱耀学院; 清华大学 人工智能产业研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出STRATA模型,以244633个参数的紧凑结构,直接从原始日内数据预测A股次日股票收益率排序,在风格残差秩IC等指标上优于参数匹配基线。

AI 中文摘要

我们提出STRATA(交错时标残差架构),这是一个拥有244633个参数的序列模型,可直接将五个交易日的原始五分钟K线和订单簿数据映射至次日的横截面收益率排序,无需人工构造特征。原始输入设置存在结构性障碍:价格序列非平稳,且不同股票间的价格水平存在数个数量级的差异,因此模型易锁定价格水平而非动态变化。STRATA通过五分支结构解决该问题:四个可学习因果深度卷积(其有效核初始化为零和),加上一个跨场线性对比;随后是四个选择性状态空间块(其衰减偏置在堆叠中交错设置)及一个四路径读出模块。由于仅偏向常见风格因子的得分在原始秩相关系数上表现良好,所有模型的得分在计算任何指标前均需针对八个量价风格因子进行残差处理。STRATA在涵盖约一千只中型市值中国A股的四年数据上训练,并在留出的一年数据上评估,其风格残差秩信息系数达0.0728(信息比率1.128,多空信号夏普比率12.85),在所有四个报告指标上均优于六个参数匹配的序列基线模型;在秩IC上,STRATA与每个基线模型的日级配对差距在p<0.001时具有统计学显著性,且在预测能力相当的模型中,STRATA的得分受控制变量的解释程度最低。收盘价目标在得分生成前已出现:改用首个可执行价格衡量时,十分位利差与零无差异,而七个架构的排序保持不变,且STRATA的优势进一步扩大。

英文摘要

We present STRATA (Staggered-Timescale Residual Architecture), a 244,633-parameter sequence model that maps five trading days of raw five-minute bar and order-book data directly to a next-day cross-sectional return ranking, with no hand-crafted features. The raw-input setting has a structural obstacle: price series are non-stationary and differ across stocks by orders of magnitude, so a model easily latches onto price level rather than dynamics. STRATA addresses it with a stem of five branches--four learnable causal depthwise convolutions whose effective kernels are initialised to sum to zero, plus one cross-field linear contrast--followed by four selective state-space blocks whose decay biases are staggered across the stack and a four-path readout. Because a score that merely tilts toward common style factors scores well on raw rank correlations, every model's scores are residualised against eight price-volume style factors before any metric is computed. Trained on four years of data covering roughly one thousand mid-capitalisation Chinese A-shares and evaluated once on a held-out year, STRATA reaches a style-residualised rank information coefficient of 0.0728 (information ratio 1.128, signal long-short Sharpe 12.85), ahead of six parameter-matched sequence baselines on all four reported metrics; on rank IC the day-level paired gap against every baseline is significant at p < 0.001, and among the arms competitive on predictive power STRATA's scores are the least explained by the controls. The close-to-close target opens before the score exists: measured instead from the first executable price, the decile spread is indistinguishable from zero, while the ordering of the seven architectures is unchanged and STRATA's margin widens.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑