发表机构
New York University(纽约大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究通过脑电数据测试视觉模型预测城市场景评价的能力,发现预测评价与神经表征对应性弱,为评估视觉模型的场景表征一致性提供了仅需55张图像嵌入的基准。
AI 中文摘要
预训练视觉嵌入正越来越多地被用作通用表征,以建模人们对城市场景的评价方式,且其验证几乎完全依赖于预测人类评分的准确性。高预测精度并不能证明这些嵌入对场景的组织方式与人类感知一致。我们针对大脑数据分别测试上述两个属性。使用公开的63名成人观看并评分56个柏林街景时的脑电图(EEG)数据,我们随时间估计场景的表征几何、该几何可解释的比例,以及其与17种特征空间的对应关系,这些特征空间涵盖语言监督、自监督、类别监督和密集预测训练,规模跨度达两个数量级,还包含可解释的对照组。对应关系整体较低:最佳表征为DINOv2 ViT-B,达到噪声下限的29.6%,各特征空间的对应范围为11.0%至29.6%,而Gabor能量描述符的表现与最佳模型无显著差异,且优于所有测试的语言监督模型。在同一模型中,更深层仍与后期神经响应匹配,因此对象识别中发现的层级对应关系,即使在整体水平较低的情况下依然存在。相同嵌入对保留的评价评分预测效果良好,相关系数r可达0.87,且两种指标在不同模型间并不相互追踪;与维度匹配的对照组相比,将特征向神经几何重新加权会降低所有测试模型的评价预测效果。因此,预测街景评价并不能作为模型以大脑方式表征街景的有力证据。该基准仅使用公开数据且无需训练,因此评估新表征仅需其对55张图像的嵌入即可。
英文摘要
Pretrained vision embeddings are widely used to model how people appraise urban scenes, and they are validated almost entirely by how well they predict human ratings. Accurate prediction shows that an embedding contains the information needed to recover the ratings. It does not show that the embedding arranges scenes as the human visual system does, which is assumed when distances or dimensions in the embedding are read as perceptual. Using openly released EEG recorded while 63 adults viewed and rated street scenes of Berlin, we compared the neural representational geometry of the scenes with the geometry of pretrained vision models that differ in training objective and size, of simple image descriptors and of the ratings themselves, each relative to a noise ceiling given by the agreement between participants. No feature space reached more than about half of the noise ceiling, which corresponds to about a fifth of the reliable variance in the neural geometry. A descriptor of oriented edge energy reached the level of most pretrained models while capturing a different part of the neural geometry, and deeper layers corresponded to later neural responses. The same embeddings predicted held-out ratings well, up to r = 0.87, yet prediction accuracy and neural correspondence were not reliably related across models, and the best predictor was among the least aligned. Predicting how a street is appraised is therefore weak evidence that a model represents the street as the brain does.