arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37230cs.CVcs.LG

看见不等于能够指称:审计冻结视觉几何中的语言可达性

Seeing Is Not Addressing: Auditing Linguistic Access to Frozen Visual Geometry

Woosang Jeon, Jiwon Yang, Soo Chung, Taehyeong Kim

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过FactorAtlas测试平台分离视觉可辨别性与语言可达性,发现匹配的视觉基础能探测并减少文本到图像检索中的访问差距。

中文摘要 AI 辅助

视觉区分往往比语言概念化所反映的更为精细。视觉语言模型表现出类似的不对称性:一个区分在冻结的图像几何中仍可辨别,但通过原生文本界面却难以指称。我们通过将文本到图像检索中的视觉可辨别性与语言可达性分离来研究这一差距。利用FactorAtlas,一个包含23,040张图像、涵盖形状、色调、图案及干扰变化的完全交叉测试平台,我们在相同区分的保留图像上比较了两种读出方式。随后,我们推导出图像侧对比,将每个值与匹配视觉基础中的替代项分离,并测试这是否减少了跨因素和模型的文本访问差距。方向特定和视觉缺失控制将这些收益归因于相关的视觉对比;这些收益在全局对齐后依然存在,并扩展到组合检索和自然图像。总之,这些结果表明视觉可辨别性与语言可达性不必一致,且匹配的视觉基础可以探测并减少由此产生的访问差距。

英文摘要

Visual distinctions are often finer than those reflected in linguistic conceptualization. Vision-language models exhibit a similar asymmetry: a distinction can remain discriminable in frozen image geometry while being weakly addressable through the native text interface. We study this gap by separating visual discriminability from linguistic addressability in text-to-image retrieval. Using FactorAtlas, a fully crossed testbed of 23,040 images spanning shape, hue, pattern, and nuisance variation, we compare both readouts on held-out images of the same distinctions. We then derive image-side contrasts that separate each value from its alternatives for matched visual grounding, and test whether this reduces the native-text access gap across factors and models. Direction-specific and visual-absence controls tie these gains to the relevant visual contrast; the gains persist after global alignment and extend to compositional retrieval and natural images. Together, these results show that visual discriminability and linguistic addressability need not coincide, and that matched visual grounding can probe and reduce the resulting access gap.

发表机构

  • Seoul National University(首尔大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑