发表机构
York University; University of Guelph(约克大学; 圭尔夫大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出深度排序与异常检测任务,结合控制图像线索和语言表达,量化视觉语言模型的深度感知能力,发现模型对深度线索利用不足且存在语言偏差。
AI 中文摘要
在本文中,我们研究视觉语言模型(VLM)的深度感知,以隔离图像深度线索的影响,并解构视觉和语言对模型性能的影响。为此,我们结合深度排序和异常检测心理物理学任务:向VLM呈现图像,其中某个对象相对于其他相同对象处于不同深度,模型必须判断异常目标相对于观察者是更近还是更远。为创建刺激,我们从模拟和真实3D场景生成2D视图,同时控制单个图像深度线索的存在,从而实现对线索级贡献的细粒度分析。通过改变指代表达的清晰度来研究语言效应。我们还引入了一种新的度量来量化视觉与语言敏感性。应用该方法,我们创建了包含37K真实和合成图像以及147K图像-问题对的异常深度(O3-D)数据集。在O3-D上评估12个开源和商业模型,结果显示深度线索利用不足,深度排序准确率在47%至56%之间,没有模型高于随机水平。同时,我们的度量揭示了答案中强烈的语言偏差。链式思维(CoT)和上下文学习(ICL)均未显著提升性能,表明仅静态图像数据可能不足以理解深度。所有代码、图像生成流程和O3-D数据集均在此https URL公开发布。
英文摘要
In this paper, we study depth perception of vision-language models (VLMs) to isolate the effects of pictorial depth cues and disentangle vision and language influences on model performance. To this end, we combine depth-ordering and odd-one-out psychophysical tasks: the VLMs are presented with images where one object is at different depth relative to other, otherwise identical, objects, and must determine whether the odd-one-out target is closer or farther to the observer. To create stimuli, we generate 2D views from simulated and real 3D scenes while controlling the presence of individual pictorial depth cues, enabling a fine-grained analysis of cue-level contributions. Language effects are examined by varying referring expression clarity. We also introduce a novel metric to quantify vision-vs-language sensitivities. Applying this methodology, we create the Odd-One-Out Depth (O3-D) dataset with 37K real and synthetic images and 147K image-question pairs. Evaluation of 12 open-source and commercial models on O3-D shows under-utilization of depth cues and depth-ordering accuracies between 47% and 56%, with no model above chance level. At the same time, our metric reveals strong linguistic bias in the answers. Neither chain-of-thought (CoT) nor in-context learning (ICL) significantly improves performance, suggesting that static image data alone may be insufficient for depth understanding. All code, the image generation pipeline, and the O3-D dataset are publicly released at https://github.com/lyiqian/o3-d.
Comments15 pages, 7 figures, accepted to ECCV 2026 (30 pages, 13 figures, supplementary materials included)