发表机构
ETH Zurich(苏黎世联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究以Qwen3-VL-4B为对象,构建含几何形状的合成数据集,探究VLMs的关系推理能力,发现其结合了真正视觉推理与基于语言线索的捷径策略。
AI 中文摘要
视觉语言模型(VLMs)在视觉推理任务中表现出色,但目前尚不清楚它们是真正理解视觉关系,还是仅仅利用语言线索或先验等捷径策略。为探究这一点,我们使用现代视觉语言模型Qwen3-VL-4B(Bai等人,2025)来解码视觉信息在不同深度的编码方式。为此,我们构建了一个包含简单几何形状的合成数据集用于受控分析,同时设计了专门测试语言线索的查询。此外,我们修改了该数据集以测试模型对视觉证据的因果依赖。结果表明,当前的视觉语言模型将真正的视觉推理与主要基于语言线索的捷径策略相结合。
英文摘要
Vision-Language Models (VLMs) achieve strong performance in visual reasoning tasks, but it remains unclear whether they understand visual relations, or simply employ shortcuts such as language cues or priors. To investigate this, we use the Qwen3-VL-4B (Bai et al., 2025), a modern VLM, to decode how visual information is encoded across depths. For this, we propose a synthetic dataset of simple geometric shapes for controlled analysis, along with queries crafted to precisely test language cues. Furthermore, the dataset is modified to test causal reliance on visual evidence. Our results show that current VLMs combine genuine visual reasoning with shortcut strategies primarily rooted in language cues.