AI 中文总结
本研究揭示预训练视觉Transformer中对象绑定的空间衰减特性,提出有限视界模型,解释大对象弱绑定、同类区分不足等行为,并初步发现视界深度差异。
AI 中文摘要
预训练的视觉Transformer能够编码两个图像块是否属于同一对象。这一IsSameObject信号可以从冻结的块嵌入中以高精度解码,这表明对象绑定仅从自监督预训练中即可涌现。我们证明,这一单一准确率数字掩盖了信号的结构。绑定是局部的:同一对象的两个块被解码为绑定的概率随它们之间的距离单调下降,并在非零底值处趋于平稳,这种下降可由具有有限长度尺度的指数函数很好地描述。这种衰减在对象大小、三类探针、ADE20K和COCO数据集以及DINO和CLIP骨干网络上均成立,这表明它是表示而非解码器的属性。将绑定视为具有有限范围的局部空间相干性,可以解释聚合分数无法解释的一系列行为:绑定在大对象上减弱,对同类不同对象的区分不如对不同类对象的区分可靠,并将对象部分与其整体分组。相比之下,在控制对象大小后,遮挡不影响绑定。我们在控制混淆因素的情况下映射了每种行为。作为初步观察,视界及其底值在DINOv2和DINOv3中组织在不同的深度,鉴于探测层数较少且两个模型之间存在混淆,我们将其报告为提示性结果。
英文摘要
Pretrained Vision Transformers encode whether two image patches belong to the same object. This IsSameObject signal is decodable from frozen patch embeddings at high accuracy, which suggests that object binding emerges from self-supervised pretraining alone. We show that this single accuracy number hides the structure of the signal. Binding is local: the probability that two patches of the same object are decoded as bound falls off monotonically with the distance between them and levels off at a nonzero floor, a falloff well described by an exponential with a finite length scale. This decay holds across object sizes, across three families of probe, on both ADE20K and COCO, and across DINO and CLIP backbones, which indicates that it is a property of the representation rather than of the decoder. Reading binding as local spatial coherence with a finite range accounts for a set of behaviors that the aggregate score leaves unexplained: binding weakens on large objects, separates distinct objects of the same class less reliably than objects of different classes, and groups object parts with their wholes. It is, by contrast, unaffected by occlusion once object size is controlled. We map each behavior with confounds controlled. As a preliminary observation, the horizon and its floor are organized at different depths in DINOv2 and DINOv3, which we report as suggestive given the small number of layers probed and the confound between the two models.
Comments14 pages, 4 figures