arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越颜色几何:评估视觉模型中类人颜色表示

Beyond Color Geometry: Evaluating Human-Like Color Representations in Vision Models

Ayan Igali, Pakizar Shamoi

arXiv 2607.13647首次发表:更新:

AI 中文总结

研究视觉模型中类人颜色表示,通过基于人类调查数据的模糊感知模型评估,测量类别边界等三个属性,发现掩码自动编码器在超越颜色几何对齐上表现最强,不同模型在颜色表示方面有差异,类人颜色基础有多个方面不能简化为单一分数。

AI 中文摘要

视觉模型看待颜色的方式与人类相同吗?现有的颜色表示评估通常将它们与诸如 CIELAB 等几何空间或离散颜色标签进行比较。这些参考标准捕获了感知距离或类别归属,但没有体现人们组织颜色的渐变方式。我们根据一个针对人类调查数据拟合的具有 86 个渐变类别的模糊感知模型来评估颜色基础。该框架可应用于任何图像编码器,并测量三个互补属性:类别边界、类别紧凑性以及超越颜色几何单独所能解释的渐变对齐。在 11 个视觉Transformer 编码器中,类别级结果大致相似,而渐变对齐差异很大。掩码自动编码器实现了最强的超越几何的对齐,其置信区间与其他编码器不重叠。分层分析进一步表明,掩码重建朝着输出保留了这种结构。在自然图像上,MAE 全局表示表面颜色,而语言监督模型相对于前景对象更强烈地编码颜色。这些结果表明,类人颜色基础有几个不同的方面,不应简化为单一分数。

英文摘要

Do vision models see colors the way humans do? Existing evaluations of color representations usually compare them with geometric spaces such as CIELAB or with discrete color labels. These references capture perceptual distance or category membership, but not the graded way in which people organize colors. We evaluate color grounding against a fuzzy perceptual model with 86 graded categories fitted to human survey data. The framework can be applied to any image encoder and measures three complementary properties: category boundaries, category compactness, and graded alignment beyond what color geometry alone can explain. Across eleven Vision Transformer encoders, the category-level results are broadly similar, whereas graded alignment differs substantially. Masked Autoencoders achieve the strongest beyond-geometry alignment, with confidence intervals that do not overlap those of the other encoders. A layer-wise analysis further shows that masked reconstruction preserves this structure toward the output. On natural images, MAE represents surface color globally, while language-supervised models encode color more strongly in relation to the foreground object. These results show that human-like color grounding has several distinct aspects that should not be reduced to a single score.

Comments6 pages, 5 figures, 2 tables submitted to 2026 Joint 14th International Conference on Soft Computing and Intelligent Systems and 27th International Symposium on Advanced Intelligent Systems (SCIS&ISIS 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑