发表机构
National University of Singapore; Center of AI Research, VinUniversity(新加坡国立大学; 文大人工智能研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究视觉语言模型空间推理难题,提出深度序数提示(DOP)方法,该方法无需训练,将单目深度转换为序数文本提示,在伪深度提供可靠排序时可改善空间推理,简单且针对性强,与其他无需训练的深度提示方法有竞争力。
AI 中文摘要
视觉语言模型(VLMs)期望能对物理空间进行推理,如判断哪个物体更近、什么在什么后面以及物体在3D中如何排列,但它们在这种空间判断上仍有困难。通常认为展示深度图能解决问题,但研究发现这可能使性能变差。深度信息虽能到达语言模型,但下游推理难以获取,渲染的伪深度图还会干扰冻结的VLMs。为此提出深度序数提示(DOP),这是一种无需训练的方法,能将单目深度转换为针对查询对象的序数文本提示。通过基准测试发现,当伪深度提供可靠对象级排序时,DOP可改善空间推理,在强原始图像模式下基本保持中立,且与最强的无需训练的深度提示方法具有竞争力,更简单且更具针对性。
英文摘要
Vision-language models (VLMs) are expected to reason about physical space -- which object is closer, what lies behind what, and how objects are arranged in 3D -- yet they still struggle with such spatial judgments. A natural remedy is to show the model a depth map, but we find that this can make performance worse. We show that depth is not absent: it reaches the language model, but becomes difficult to access for downstream reasoning, while rendered pseudo-depth maps act as noisy auxiliary images that frozen VLMs cannot easily regulate. We propose Depth-Ordinal Prompting (DOP), a training-free method that converts monocular depth into a single question-targeted ordinal text cue at the queried objects, without adding a depth image, training a module, injecting features, or using labels. Our key finding is form dependence: the same depth signal can hurt when shown as an image but help when told as text.Across benchmarks, models, and depth estimators, DOP improves spatial reasoning when pseudo-depth provides reliable object-level ordering and remains largely neutral in strong original-image regimes. It is also competitive with the strongest training-free depth-prompting alternative while being simpler and more targeted.
CommentsWork in progress