发表机构
Computer Vision Center; Universitat Autònoma de Barcelona; Image Processing Lab, Universitat de València; Arab Academy for Science, Technology, and Maritime Transport(计算机视觉中心; 巴塞罗那自治大学; 瓦伦西亚大学图像处理实验室; 阿拉伯科学、技术和海运学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究探索50多个预训练视觉编码器的色度敏感性,将其与人类辨别阈值比较,发现模型表示与人类感知阈值对齐弱,自监督编码器优于监督的,语言监督模型两极分化,表明当前训练目标不会自然产生类人色度敏感性。
AI 中文摘要
理解和表征人类颜色感知是一个长期的研究目标。最传统的方法之一是寻找人类颜色辨别阈值,即人类观察者可感知的最小色度差异。近年来,深度神经网络已成为计算机视觉任务的标准网络。特别是深度视觉编码器,在大规模视觉数据上训练的基础模型,将图像映射到潜在特征表示。尽管深度视觉编码器被广泛使用,但很少有研究调查其内部表示是否表现出类人的辨别阈值。在这项工作中,我们进行了一项大规模探索性研究,探究50多个预训练视觉编码器(包括卷积网络和视觉Transformer)对人类辨别阈值的色度敏感性。使用多个色度水平的受控色度刺激,我们通过区域重叠度量(mIoU)将模型得出的色度辨别阈值与人类辨别椭圆进行比较。我们的分析表明,所有模型家族的模型表示与人类感知阈值之间的对齐通常较弱,最佳mIoU < 0.25。此外,我们发现自监督编码器始终优于监督编码器,而语言监督模型表现出最两极分化的行为,在排名中占据顶部和底部。这些发现表明,对于任何分析的架构,类人的色度敏感性都不会从当前的大规模视觉训练目标中自然出现。
英文摘要
Understanding and characterizing human color perception is a longstanding research goal. One of the most traditional approaches is looking for the human color discrimination thresholds, the minimum chromatic differences perceptible to human observers. In recent years, deep neural networks have become the standard networks for computer vision tasks. In particular, deep vision encoders, foundation models trained on large-scale visual data, map images into latent feature representations. Despite the widespread use of deep vision encoders, few studies have investigated whether their internal representations exhibit human-like discrimination thresholds. In this work, we present a large-scale exploratory study probing the chromatic sensitivity of more than 50 pretrained vision encoders, including convolutional networks and vision transformers, against human discrimination thresholds. Using controlled chromatic stimuli at multiple chroma levels, we compare model-derived chromatic discrimination thresholds with human discrimination ellipses through a region-overlap metric (mIoU). Our analysis reveals generally weak alignment between model representations and human perceptual thresholds across all model families, with the best mIoU < 0.25. Moreover, we find that self-supervised encoders consistently outperform supervised ones, while language-supervised models show the most polarized behavior, occupying both the top and bottom of the ranking. These findings suggest that human-like chromatic sensitivity does not emerge naturally from current large-scale visual training objectives for any of the analyzed architectures.