发表机构
The University of Tokyo; Google DeepMind(东京大学; 谷歌DeepMind)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过关系匹配到样本范式与机制分析,揭示VLM抽象推理由早期对象特征与晚期关系两个竞争回路实现,并发现能力层级、规模等四杠杆影响关系匹配。
AI 中文摘要
视觉语言模型(VLMs)在视觉基准测试中表现出色,但在需要抽象推理的任务上却系统性失败。现有基准测试记录了这一失败,但无法说明其发生的原因或缺失了哪种认知能力。我们通过采用比较心理学和发展心理学中的关系匹配到样本(RMTS)范式,并结合对模型内部的机制分析来弥合这一差距。在一个参数化控制的刺激集上,跨越前沿API模型(GPT、Claude、Gemini)和三个开源系列(Qwen3.5、Gemma-4、InternVL3)进行评估,我们确定了四个促使VLM转向关系匹配的杠杆——能力层级、模型规模、每个场景中的对象数量以及每对象刺激噪声的缺失——共同产生了一条类似发展的轨迹,反映了人类的“关系转变”。打开模型,逐层表征相似性分析和因果中介分析揭示,VLM的抽象推理由两个竞争回路实现:一个早期回路根据表面对象特征组织图像,一个晚期回路根据抽象关系组织图像。将分析扩展到ARC-AGI-1,我们发现消融在RMTS上识别出的关系头比消融随机头更能降低性能,表明关系回路在我们的受控刺激之外也被调用。我们希望这种机制层面的视角能成为理解抽象推理如何在VLM中实现的一步。
英文摘要
Vision Language Models (VLMs) excel on visual benchmarks but fail systematically on tasks requiring abstract reasoning. Existing benchmarks document this failure but cannot say \emph{why} it happens or which cognitive capability is missing. We close this gap by adopting the Relational Match-to-Sample (RMTS) paradigm from comparative and developmental psychology and pairing it with a mechanistic analysis of the model's internals. On a parametrically controlled stimulus set evaluated across frontier API models (GPT, Claude, Gemini) and three open-source families (Qwen3.5, Gemma-4, InternVL3), we identify four levers that shift VLMs toward the relational match---capability tier, model scale, the number of objects per scene, and the absence of per-object stimulus noise---together producing a developmental-like trajectory that mirrors the human \emph{relational shift}. Opening up the model, a per-layer representational similarity analysis and a causal mediation analysis reveal that VLM abstract reasoning is implemented by two competing circuits: an early circuit that organises images by their surface object features, and a late circuit that organises them by their abstract relation. Extending the analysis to ARC-AGI-1, we find that ablating the relational heads identified on RMTS degrades performance more than ablating random heads, indicating that the relational circuit is recruited beyond our controlled stimuli. We hope this mechanism-level view serves as a step toward understanding how abstract reasoning is implemented in VLMs.