AI 中文总结
针对EEG基础模型存在的低频偏差问题,提出频率平衡掩码自编码框架FAME,在OmniEEG-Bench的41项下游任务中24项达到SOTA,凸显平衡频谱监督的重要性。
AI 中文摘要
增大脑电(EEG)预训练数据规模或模型容量并不能持续提升下游任务性能。我们发现各类EEG基础模型学习到的表征中存在持续的低频偏差,该偏差在不同数据集规模、模型容量和预训练目标下均存在。我们的分析将该偏差与EEG类1/f^α的频谱结构和神经网络优先学习低频分量的倾向之间的关联起来。在掩码自编码器中,ℓ₂重建目标进一步放大了这种不平衡:在相当的相对重建误差下,高功率低频分量对损失的贡献不成比例。为解决该问题,我们提出FAME,一种频率平衡的掩码自编码框架,该框架从掩码EEG输入中重建预定义EEG频带内的时频活动。FAME对每个频带内的重建目标进行独立标准化,并为所有频带特定损失分配相等权重,从而平衡EEG频谱上的监督。在OmniEEG-Bench的41项下游任务上评估显示,FAME学习到了更具频谱平衡性的表征,并在其中24项任务上取得了SOTA性能。这些结果强调了平衡频谱监督对于学习可迁移EEG表征的重要性。
英文摘要
Increasing EEG pretraining data scale or model capacity does not consistently improve downstream performance. We identify a persistent low-frequency bias in representations learned by diverse EEG foundation models, which remains across dataset scales, model capacities, and pretraining objectives. Our analysis links this bias to the interaction between EEG's $1/f^α$-like spectral structure and neural networks' tendency to preferentially learn low-frequency components. In masked autoencoders, the $\ell_2$ reconstruction objective further amplifies this imbalance: under comparable relative reconstruction errors, high-power low-frequency components contribute disproportionately to the loss. To address this issue, we introduce FAME, a frequency-balanced masked autoencoding framework that reconstructs time--frequency activity in predefined EEG bands from masked EEG inputs. FAME independently standardizes the reconstruction targets within each band and assigns equal weight to all band-specific losses, thereby balancing supervision across the EEG spectrum. Evaluated on 41 downstream tasks in OmniEEG-Bench, FAME learns more spectrally balanced representations and achieves state-of-the-art performance on 24 of them. These results underscore the importance of balanced spectral supervision for learning transferable EEG representations.