发表机构
Centre for Human-Computer Interaction, Indian Institute of Technology; Freie Universität; Goethe University(印度理工学院人机交互中心; 柏林自由大学; 法兰克福大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究比较CNN与视觉Transformer在预测人类EEG响应上的表现,发现Transformer在深层保持更强的脑预测对应关系,且保留补丁令牌是关键,架构而非训练目标驱动此差异。
AI 中文摘要
卷积神经网络(CNNs)和视觉Transformer都被用于建模人类视觉系统,但两种架构是否在网络深度的特定点上产生分歧尚不清楚。我们通过计算每个模型在每一层或每个块上的预测EEG响应与实测EEG响应之间的皮尔逊相关系数(r),比较了六个CNN和两个视觉Transformer,实验中有十名参与者观看了200张自然图像。对于Transformer模型,我们还测试了四种令牌表示,从仅使用分类(CLS)令牌到CLS与所有补丁令牌的组合。CNN在最早层表现出最强的对应关系,在更深层逐渐减弱,尤其是在刺激后响应的后期。相反,Transformer在其最深的块中保持了较强的对应关系,尽管在其最早的块中并非如此。这种优势取决于令牌表示:池化表示给出的峰值相关性较弱(r约0.48-0.51),而保留所有补丁令牌的表示则更强(CLIP-ViT-B/32的r=0.640,DINOv2-ViT-B/14的r=0.656)。受控比较表明,是架构而非训练目标驱动了这种效应:MoCo-v1和ResNet-50(架构匹配)的表现几乎相同(r=0.673, 0.670),而CLIP-RN50和CLIP-ViT-B/32(目标匹配)在保留补丁令牌之前出现分歧。我们提出,CNN训练的分类瓶颈在深度上压缩了与大脑相关的信息,而Transformer的自注意力机制和非分类目标则不然。空间地形图分析显示,所有模型均呈现常见的枕叶主导模式,表明这些差异反映了信号强度和持续性,而非不同的脑区。保留补丁的Transformer表示在CNN崩溃的地方维持了大脑预测的对应关系。
英文摘要
Convolutional neural networks (CNNs) and vision transformers are both used to model the human visual system, but whether the two architectures diverge at a specific point in network depth is unclear. We compared six CNNs and two vision transformers by computing the Pearson correlation (r) between each model's predicted and measured EEG response at every layer or block, in ten participants viewing 200 natural images. For the transformer models, we also tested four token representations, from the classification (CLS) token alone to CLS combined with all patch tokens. CNNs showed strongest correspondence at the earliest layers, weakening at deeper layers, particularly later in the post-stimulus response. Transformers instead sustained strong correspondence at their deepest blocks, though not at their earliest ones. This advantage depended on token representation: pooled representations gave weaker peak correlations (r approx 0.48-0.51) than representations retaining all patch tokens (r=0.640 for CLIP-ViT-B/32, r=0.656 for DINOv2-ViT-B/14). Controlled comparisons showed architecture, not training objective, drove this effect: MoCo-v1 and ResNet-50 (matched architecture) performed nearly identically (r=0.673, 0.670), whereas CLIP-RN50 and CLIP-ViT-B/32 (matched objective) diverged until patch tokens were preserved. We propose that CNN training's classification bottleneck compresses brain-relevant information at depth, unlike transformers' self-attention and non-classification objectives. A spatial topography analysis showed a common occipital-dominant pattern across all models, indicating these differences reflect signal strength and persistence rather than distinct brain regions. Patch-preserving transformer representations sustain brain-predictive correspondence where CNNs collapse.