arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SPD-MetaFormer是小样本脑解码的理想选择

SPD-MetaFormer is what you need for small-data brain decoding

Zhida Wang, Wei Lyu, Guo Yu, Sui Tang

arXiv 2610.10952首次发表:更新:

AI 中文总结

针对小样本脑解码问题,本文提出无注意力的SPD-MetaFormer架构,在三个EEG基准测试中取得与现有方法相当的性能,为自适应流形注意力提供了更简单的替代方案。

AI 中文摘要

脑信号解码面临挑战,原因在于神经记录存在噪声且个体间存在差异,同时带标签数据往往有限。近期基于对称正定(SPD)流形的注意力模型利用协方差和连通性表示取得了良好性能,但学习到的令牌权重的作用仍不明确。我们研究了两种代表性架构:基于对数欧几里得几何的MAtt和基于广义Bures-Wasserstein几何的GBWAtt,发现它们训练后学习到的注意力权重仍接近均匀分布。我们将此行为归因于有界相似性参数化,在原始softmax缩放下,该参数化限制了注意力权重的对比度。此外,在整个训练和评估过程中,用均匀权重替代学习到的权重,对平均预测性能几乎没有影响,同时保留了每个模型的原始聚合几何。受这些发现的启发,我们引入了SPD-MetaFormer,这是一种无注意力的架构,基于对数欧几里得几何下的均匀加权Fréchet聚合构建。其主干使用测地残差更新摘要令牌,并使用共享谱前馈映射变换所有令牌,随后进行学习到的加权读出。令牌状态保持为SPD值,直到切空间分类。在三个EEG基准测试中,SPD-MetaFormer相对于已发表的欧几里得和流形基线取得了有竞争力的结果。单独的匹配复现测试了MAtt和GBWAtt中学习到的权重与均匀权重的效果。这些结果表明,在所研究的短序列和有限数据场景下,精心设计的SPD架构可以为自适应流形注意力提供更简单且有效的替代方案。

英文摘要

Brain signal decoding is challenging because neural recordings are noisy and vary across individuals, while labeled data are often limited. Recent attention-based models on the symmetric positive definite (SPD) manifold have nevertheless achieved strong performance using covariance and connectivity representations, yet the contribution of learned token weighting remains unclear. We examine two representative architectures, MAtt (based on log-Euclidean geometry) and GBWAtt (based on generalized Bures--Wasserstein geometry), and find that their learned attention weights remain close to uniform after training. We relate this behavior to bounded similarity parameterizations that, under the original softmax scaling, limit attention-weight contrast. Moreover, replacing learned weights with uniform weights, throughout training and evaluation, has little effect on mean predictive performance while preserving each model's original aggregation geometry. Motivated by these findings, we introduce SPD-MetaFormer, an attention-free architecture built on uniformly weighted Fréchet aggregation under log-Euclidean geometry. Its backbone uses a geodesic residual to update a summary token and a shared spectral feedforward map to transform all tokens, followed by a learned weighted readout. Token states remain SPD-valued until tangent-space classification. Across three EEG benchmarks, SPD-MetaFormer achieves competitive results relative to published Euclidean and manifold baselines. Separate matched reproductions test learned versus uniform weighting within MAtt and GBWAtt. These results suggest that, in the short-sequence and limited-data regimes studied, carefully designed SPD architectures can provide a simpler and effective alternative to adaptive manifold attention.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑