arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00726cs.CV

凹点探测法在视觉基础模型中恢复局部绑定信息

Foveated Probes Recover Localized Binding Information in Vision Foundation Models

Mateusz Michalkiewicz, Mahsa Baktashmotlagh, Guha Balakrishnan

首次发表
浏览论文内容

中文总结 AI 辅助

该研究通过对比全局读出、凹点读出和神谕读出,发现冻结视觉基础模型的空间盲性源于全局嵌入界面,而非 patch 标记缺乏空间信息,凹点读出可有效恢复局部绑定信息。

中文摘要 AI 辅助

冻结的视觉基础模型通常通过单个全局图像嵌入进行评估,但这种界面可能会将信息缺失与读出阶段丢失的信息混为一谈。我们通过保持预训练视觉编码器冻结,仅改变应用于其最终 patch 标记的读出方式来研究这种区别。我们将标准全局读出与轻量型凹点读出(使用学习到的或问题条件化的查询对 patch 标记进行注意力池化),以及可访问标注目标区域的神谕读出进行对比。我们在三个局部绑定问题上评估这些界面:杂乱环境下的受控合成颜色-形状绑定任务、无颜色的拥挤形状检测变体,以及 GQA 衍生的自然图像任务(其中配对问题要求查询同一图像中不同同类对象的颜色)。当合成目标单独出现时,全局读出表现近乎完美,但在杂乱环境和反事实目标编辑下会失效;而凹点读出可恢复大部分神谕可访问的信号。在 GQA 衍生任务中,与仅基于问题的先验相比,不依赖问题的全局图像向量仅略有提升,而问题条件化的凹点机制则大幅提高了配对局部颜色的准确率。反事实干扰-信号比解释了合成任务中的失效:全局池化会稀释改变局部标签的证据,同时使探测器暴露于无关对象的干扰变化中。这些结果表明,冻结视觉模型中明显的空间盲性可能源于全局嵌入界面,而非冻结 patch 标记中缺乏空间信息。

英文摘要

Frozen vision foundation models are commonly evaluated through a single global image embedding, but this interface can conflate missing information with information lost at readout time. We study this distinction by keeping a pretrained vision encoder frozen and varying only the readout applied to its final patch tokens. We compare standard global readouts against a lightweight foveated readout, which attention-pools patch tokens using a learned or question-conditioned query, and against an oracle readout with access to the annotated target region. We evaluate these interfaces on three localized binding problems: a controlled synthetic color--shape binding task under clutter, a color-free crowded shape-detection variant, and a GQA-derived natural-image task where paired questions ask for the colors of different same-category objects in the same image. Global readouts perform near perfectly when the synthetic target appears alone, but collapse under clutter and counterfactual target edits, whereas the foveated readout recovers most of the oracle-accessible signal. On the GQA-derived task, question-independent global image vectors improve only modestly over question-only priors, while question-conditioned foveation substantially improves paired localized color accuracy. A counterfactual nuisance-to-signal ratio explains the synthetic failures: global pooling dilutes localized label-changing evidence while exposing the probe to nuisance variation from irrelevant objects. These results indicate that apparent spatial blindness in frozen vision models can arise from the global embedding interface rather than from an absence of spatial information in the frozen patch tokens.

补充信息

↑