AI 中文总结
针对M/EEG视觉解码中全局嵌入对齐的局限,提出残差补丁适配器RPA,利用ViT中间层全部补丁令牌进行对齐,在THINGS-EEG2和THINGS-MEG上达到SOTA性能,并验证了高层语义与低层特征的重要性。
AI 中文摘要
现有的大多数脑磁图和脑电图(M/EEG)视觉解码方法将大脑信号与从预训练视觉编码器中提取的单一全局嵌入进行对齐,但中间补丁表示(保留了更丰富、更细致的视觉信息)是否能够改善表示学习,这一问题仍未得到解答。为解决此问题,我们引入了残差补丁适配器(RPA),这是一种轻量级、模块化的适配器,利用ViT视觉编码器中间层的所有补丁令牌进行对齐。通过广泛的消融分析,我们首先表明,对补丁令牌进行池化或掩蔽会降低学习到的表示质量,证明保留完整的补丁令牌集对于脑电图对齐至关重要,而CLS令牌提供的信息独特且有限。随后,我们使用一系列六项定量特征分析表明,高层语义和低层视觉特征(包括颜色和纹理)对于这种脑电图到图像的比对均不可或缺。在当前协议下,我们的系统在THINGS-EEG2数据集上实现了受试者内Top-1准确率95.4%和跨受试者35.5%,在THINGS-MEG数据集上分别达到65.2%和6.7%,在两个数据集上均实现了最先进的(SOTA)性能。使用替代性脑编码器(包括预训练的脑电图基础模型)进行的评估表明,该方法不仅限于基于投影的脑电图编码器。此外,我们提供了一个即插即用接口,允许用卷积、注意力或ConvNeXt替代方案替换RPA。综合这些发现,通过建立利用视觉编码器潜在空间的设计原则,为M/EEG到图像的表示学习提供了重要见解,并为大脑-图像对齐和非侵入式脑机接口(BCI)开辟了新方向。
英文摘要
Most existing MEG and EEG (M/EEG) visual decoding methods align brain signals with a single global embedding extracted from a pretrained visual encoder, leaving open whether intermediate patch representations, which preserve richer and more granular rich visual information, can improve representation learning. To address this question, we introduce the Residual Patch Adapter (RPA), a lightweight, modular adapter that leverages all patch tokens from an intermediate layer of a ViT visual encoder for alignment. Through extensive ablation analyses, we first show that pooling or masking patch tokens degrades the learned representation, demonstrating that retaining the full set of patch tokens is important for EEG alignment, while the CLS token provides little unique information. We then use a series of six quantitative feature analyses to show that both higher-level semantics and lower-level visual features, including color and texture, are essential for this EEG-to-image alignment. Under current protocols, our system achieves Top-1 accuracies of 95.4\% within-subject and 35.5\% cross-subject on THINGS-EEG2, and 65.2\% and 6.7\%, respectively, on THINGS-MEG, achieving state-of-the-art (SOTA) performance across both datasets. Evaluations with alternative brain encoders, including pretrained EEG foundation models, demonstrate that the approach extends beyond the projection-based EEG encoder. Furthermore, we provide a plug-and-play interface that allows RPA to be replaced by convolution, attention, or ConvNeXt alternatives. Together, these findings provide significant insight into M/EEG-to-image representation learning by establishing design principles for leveraging the latent space of visual encoders, and open new directions for brain--image alignment and non-invasive brain--computer interface (BCI).
Comments34 pages, 13 figures