Nyström 注意力在横截面股票预测中与全注意力相当
Nyström Attention Matches Full Attention for Cross-Sectional Stock Prediction
- The University of Hong Kong(香港大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文分解 MASTER 的股票间注意力,发现其近似均匀但低秩,Nyström 低秩近似可匹配全注意力性能,且注意力寻求互补而非相关性,大规模下跨股票模块无优势。
AI中文摘要:
MASTER 的股票间多头注意力——负责建模横截面股票关系的模块——占模型参数的 42.5% 和预测价值的 25%。我们系统地分解该模块,并揭示了一个令人惊讶的结构:学习到的注意力近似均匀(困惑度 278/300),然而强制完全均匀会消除所有横截面判别能力。频谱分析解决了这一悖论:与均匀分布的偏差是低秩的(有效秩约 65,前 10 个模态捕获 96.5% 的能量),这解释了为什么稀疏近似持续失败,而 Nyström 低秩注意力(m=32 个地标)在 O(mN) 成本下匹配全 O(N^2) 注意力——通过 TOST 在 N=300(5 个种子,Rank IC p=0.003)和 N=800(10 个种子,Rank IC p=0.034)下认证等价。额外发现包括:(i)注意力与收益相似性负相关(Spearman rho = -0.614;在行业标记子集上,无条件为 -0.645,在控制行业、beta 和波动性后为 -0.627),表明寻求互补而非相关性挖掘;(ii)所有基于图的替代方案都会降低性能,硬掩蔽比完全移除模块更差;(iii)在 N ~ 3,500 且采用适应架构时,没有跨股票模块(GCN、Nyström 或 MASTER 风格流水线)显著优于每股票 LSTM 基线(n=4 个种子),表明在较小规模下观察到的收益不能简单迁移。这些结果确立了股票间注意力的价值在于一种可压缩、动态、近全局的重分配,这种重分配奖励低秩近似但抵抗稀疏化。
英文摘要:
MASTER's inter-stock multi-head attention -- the module responsible for modeling cross-sectional stock relationships -- accounts for 42.5% of model parameters and 25% of predictive value. We systematically decompose this module and uncover a surprising structure: the learned attention is near-uniform (perplexity 278/300), yet forcing exact uniformity eliminates all cross-sectional discrimination. Spectral analysis resolves this paradox: the deviation from uniformity is low-rank (effective rank ~65, top-10 modes capture 96.5% of energy), explaining why sparse approximations consistently fail while Nystrom low-rank attention (m=32 landmarks) matches full O(N^2) attention at O(mN) cost -- certified equivalent via TOST at both N=300 (5 seeds, Rank IC p=0.003) and N=800 (10 seeds, Rank IC p=0.034). Additional findings include: (i) attention anti-correlates with return similarity (Spearman rho = -0.614; on the industry-labeled subset, -0.645 unconditionally and -0.627 after controlling for industry, beta, and volatility), suggesting complementarity-seeking rather than correlation mining; (ii) all graph-based alternatives degrade performance, with hard masking worse than complete module removal; and (iii) at N ~ 3,500 with adapted architectures, no cross-stock module (GCN, Nystrom, or MASTER-style pipeline) significantly outperforms a per-stock LSTM baseline (n=4 seeds), indicating that the benefits observed at smaller scales do not trivially transfer. These results establish that the inter-stock attention's value resides in a compressible, dynamic, near-global redistribution that rewards low-rank approximation but resists sparsification.