arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉Transformer从自然图像中学习格式塔样的图形-背景线索

Vision Transformers Learn Gestalt-Like Figure-Ground Cues from Natural Images

Matthias Tangemann, Benjamin Lo, Zygmunt Pizlo, Kaleem Siddiqi, Dirk B. Walther, Sven Dickinson

arXiv 2607.08932首次发表:更新:

发表机构

University of Toronto; Vector Institute; McGill University; MILA; UC Irvine(多伦多大学; 向量研究所; 麦吉尔大学; 蒙特利尔学习算法研究院; 加州大学欧文分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究评估视觉Transformer中基于形状的图形-背景组织,通过拟合线性探针,用自然图像和人工刺激测试25个ViT,发现其能编码被包围性和凸性,自然图像训练的探针可零样本泛化,对称性结果有别,证明线索可从自然场景学,ViT是研究感知组织的有力模型。

AI 中文摘要

人类视觉系统中的图形-背景组织依赖于多种基于形状的线索,包括被包围性、凸性和对称性。虽然这些线索已通过抽象刺激得到广泛研究,但对于它们在自然条件下如何运作或如何从自然场景统计中产生却知之甚少。深度神经网络提供了一条有前景的道路:一个依赖与人类相同图形-背景线索的模型将为潜在机制提供易于处理的实验途径。在本研究中,我们评估了视觉Transformer(ViT)中基于形状的图形-背景组织,此前工作已证明其出现基于对象的分组。我们通过拟合线性探针,使用自然图像和隔离单个线索的受控人工刺激,从中间补丁表示预测图形-背景分配,测试了25个跨越监督和自监督训练目标的ViT。结果表明,ViT能稳健地编码被包围性和凸性,且在自然图像上训练的探针能在多个模型上零样本泛化到人工刺激。对于对称性,结果不一:均匀着色区域能编码该线索,纹理区域则不能。总之,我们的发现表明格式塔样的图形-背景线索可从自然场景统计中学习,并将ViT定位为研究感知组织计算机制的有力模型系统。代码和数据可在该https网址获取。

英文摘要

Figure-ground organization in the human visual system relies on several shape-based cues, including surroundedness, convexity, and symmetry. While these cues have been extensively studied using abstract stimuli, little is known about how they operate under natural conditions or how they arise from the statistics of natural scenes. Deep neural networks offer a promising path forward: a model that relies on the same figure-ground cues as humans would provide tractable experimental access to the underlying mechanisms. In this study, we evaluate shape-based figure-ground organization in Vision Transformers (ViTs), for which prior work has demonstrated the emergence of object-based grouping. We test 25 ViTs spanning supervised and self-supervised training objectives, by fitting linear probes to predict figure-ground assignment from intermediate patch representations using both natural images and controlled artificial stimuli that isolate individual cues. Our results show that ViTs robustly encode surroundedness and convexity, and that probes trained on natural images generalize zero-shot to artificial stimuli across several models. For symmetry we observe mixed results: the cue is encoded for uniformly colored but not for textured regions. Taken together, our findings demonstrate that Gestalt-like figure-ground cues can be learned from natural scene statistics and position ViTs as a compelling model system for studying the computational mechanisms of perceptual organization. Code and data is available at https://github.com/mtangemann/mlvbench

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑