符号 grounding 揭示了抽象视觉推理中的表征瓶颈
Symbolic Grounding Reveals Representational Bottlenecks in Abstract Visual Reasoning
- Applied Artificial Intelligence Group, Tallinn University of Technology, Estonia(塔林技术大学应用人工智能小组,爱沙尼亚)
- Kimova AI, Tallinn, Estonia(Kimova AI,塔林,爱沙尼亚)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文通过Bongard-LOGO基准测试,发现符号输入能显著提升抽象视觉推理性能,揭示了表征能力是核心瓶颈。
AI中文摘要:
视觉--语言模型(VLMs)在抽象视觉推理基准测试中(如Bongard问题)往往表现不佳,这引发了关于瓶颈在于推理还是表征的疑问。我们通过Bongard-LOGO基准测试研究这一问题,Bongard-LOGO是一个具有真实生成程序的合成抽象概念学习基准测试,通过比较端到端的VLMs在原始图像上的表现与大型语言模型(LLMs)在基于这些图像的符号输入下的表现进行比较。使用符号输入作为诊断探针而非实用多模态架构,我们的组件--语法(C--G)范式将Bongard-LOGO重新表述为基于LOGO风格动作程序或结构化描述的符号推理任务。LLMs在抽象推理任务上取得了显著且一致的提升,达到自由形式问题的中90年代准确率,而强大的视觉基线在匹配的任务定义下仍接近随机水平。对输入格式、显式概念提示和最小视觉表征的消融实验表明,这些因素远不如从像素到符号结构的转变重要。这些结果将表征确定为抽象视觉推理中的关键瓶颈,并展示了符号输入如何作为受控诊断上限。
英文摘要:
Vision--language models (VLMs) often fail on abstract visual reasoning benchmarks such as Bongard problems, raising the question of whether the main bottleneck lies in reasoning or representation. We study this on Bongard-LOGO, a synthetic benchmark of abstract concept learning with ground-truth generative programs, by comparing end-to-end VLMs on raw images with large language models (LLMs) given symbolic inputs derived from those images. Using symbolic inputs as a diagnostic probe rather than a practical multimodal architecture, our \emph{Componential--Grammatical (C--G)} paradigm reformulates Bongard-LOGO as a symbolic reasoning task based on LOGO-style action programs or structured descriptions. LLMs achieve large and consistent gains, reaching mid--90s accuracy on Free-form problems, while a strong visual baseline remains near chance under matched task definitions. Ablations on input format, explicit concept prompts, and minimal visual grounding show that these factors matter much less than the shift from pixels to symbolic structure. These results identify representation as a key bottleneck in abstract visual reasoning and show how symbolic input can serve as a controlled diagnostic upper bound.