AI 中文总结
研究发现早期视觉皮层的模型-大脑RSA中,未训练与反向传播训练网络的表现差异随评估分辨率变化,该效应由图像细节承载,且限定了此类比较方法的分辨能力。
AI 中文摘要
表征相似性分析(RSA)正越来越多地用于探究哪些学习规则能赋予卷积神经网络类脑表征。由于反馈对齐、预测编码、STDP等生物可行的学习规则无法扩展,包含这些规则的研究通常在小图像(通常为32×32的CIFAR数据集)上训练小型网络,之后将其与高得多分辨率下建模的大脑响应进行比较。我们发现,在此设定下一个常见结果——未训练或局部训练的网络在早期视觉皮层上的表现可与反向传播训练的网络匹敌甚至超过后者——高度依赖于评估网络的分辨率。未训练网络与反向传播训练网络在V1(初级视皮层)的差异,从32px训练分辨率下的-0.001±0.007,在224px分辨率下扩大至+0.044±0.006,且在6种分辨率(n=5个随机种子)间单调增长。该结果在人类fMRI(功能磁共振成像)中成立,在沿训练轨迹的单种子猕猴电生理学中也呈相同趋势,且适用于在224px分辨率下训练的ImageNet ResNet-50和Swin-Tiny transformer。我们测试了4种候选机制,均无法解释该效应:训练/评估分辨率匹配、低水平Gabor和像素结构、未训练基线的归一化状态、池化描述符向全局亮度统计量收敛;其中3种通过保持卷积权重完全相同的干预被排除。第五个实验定位了该效应:将图像细节限制在训练分辨率,同时让池化位置增长12倍,可消除约90%的效应,因此该依赖由图像细节而非池化承载。此外,每张图像的单个标量亮度值与V1 RDM(表征差异矩阵)的相关系数ρ=0.075,与未训练网络的0.076基本匹配,这限定了此类比较方法能分辨的范围。跨分辨率保持一致的唯一学习效应是在LOC(外侧枕叶皮层)中反向传播训练网络优于未训练网络。
英文摘要
Studies of biologically plausible learning rules (feedback alignment, predictive coding, STDP) train small networks on $32\times32$ images and compare them with brain responses to naturalistic stimuli presented at much higher resolution. We find that a common result in this setting, that untrained or locally trained networks rival or beat backpropagation at early visual cortex, depends on the resolution at which networks are evaluated. In representational similarity analysis (RSA) against human V1, the gap between an untrained and a backprop-trained network grows monotonically from $-0.001$ (95% CI $[-0.011, 0.009]$) at the 32 px training resolution to $+0.030$ $[0.018, 0.042]$ at 224 px (six resolutions, 5 seeds; per-subject RSA on cross-run stimulus pairs). The growth reproduces for three further learning rules, directionally in single-seed macaque electrophysiology, along training, and for an ImageNet ResNet-50 and a Swin-Tiny transformer, which also align best at low resolution although trained at 224 px. Gabor and pixel structure, the normalization state of the untrained baseline, and a global brightness statistic do not account for it. Against a criterion fixed before the analysis, limiting image content to 32 px while the input still grows does not shrink the V1 gap ($+0.035$ $[0.023, 0.046]$ vs. $+0.030$ at 224 px; ratio 1.14 $[1.03, 1.34]$): it arises once the input exceeds the training resolution, through a decline of backprop. A single luminance value per image reaches $ρ= 0.054$ against V1, matching the untrained network (0.053), which bounds what RSA on globally pooled features can resolve at V1 in this dataset. The one learning effect that holds across resolution is backprop above untrained at LOC. Comparisons of learning rules or architectures at early visual cortex need to control, and report, the evaluation resolution.
Commentsv3: RSA per subject on cross-run stimulus pairs; V1 gap +0.044 -> +0.030 [0.018, 0.042], still ~0 at 32 px. Leave-one-subject-out lower bound replaces the withdrawn noise ceiling. Sec. 3.5 reversed: with content limited to 32 px the V1 gap is as large (R = 1.14), so it follows input size, not detail above 32 px. Figures recomputed. Main result holds