arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

规范颜色作为视觉编码器与VLM中概念可解码性的透镜

Canonical Color as a Lens into Concept Decodability in Vision Encoders and VLMs

Xiaofu Chen, Stella Frank, Yova Kementchedjhieva

arXiv 2609.09124首次发表:更新:

发表机构

MBZUAI; Technical University of Denmark(穆罕默德·本·扎耶德人工智能大学; 丹麦技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究利用规范颜色作为受控测试案例,通过构建数据集和探测实验,发现视觉编码器能从灰度图像中解码规范颜色且与物体身份关联,并揭示VLM后训练显著影响颜色可解码性,为追踪概念语义信息提供新透镜。

AI 中文摘要

视觉编码器为视觉-语言模型构建图像输入的表示。这种表示包含多少概念性(而非直接可见的)信息?我们使用规范颜色作为受控测试案例,探究视觉编码器是否使规范颜色信息线性可访问,即使颜色已从输入图像中移除。我们构建了一个具有规范颜色的物体数据集,并使用彩色和灰度图像探测视觉编码器的颜色和物体身份信息。我们发现,规范颜色仍可从灰度图像中解码,并与预测的物体身份相关联,表明存在概念性联系。将此分析扩展到完整的VLM,我们发现VLM的后训练对视觉编码器中的颜色可解码性可能产生惊人的大影响。总体而言,规范颜色为追踪视觉编码器和VLM中物体级概念语义信息提供了一个有用的可控透镜。

英文摘要

Visual encoders construct a representation of the image input for Vision-Language models. How much conceptual, as opposed to immediately visible, information does this representation contain? We use canonical color as a controlled test case to ask whether vision encoders make canonical-color information linearly accessible, even when color is removed from the input image. We construct a dataset of objects with canonical colors, and probe vision encoders for both color and object identity using color and grayscale images. We find that canonical color remains decodable from grayscale images, and is tied to predicted object identity, indicating a conceptual link. Extending this analysis to full VLMs, we find that VLM post-training can have a surprisingly large effect on color decodability in the vision encoder. Overall, canonical color provides a usefully controllable lens for tracing object-level conceptual semantic information in vision encoders and VLMs.

Comments12 pages, 7 figures. Accepted to EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑