arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2603.23867cs.LGcs.AIcs.CV

VLM能稳健推理吗?一项神经符号研究

Can VLMs Reason Robustly? A Neuro-Symbolic Investigation

  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
  • University of Edinburgh(爱丁堡大学)

机构由 AI 辅助整理,请以论文原文为准。

Weixin Chen, Antonio Vergari, Han Zhao

AI总结:

研究视觉语言模型在分布偏移下的推理稳健性,提出结合VLM概念识别与电路符号推理的神经符号方法VLC,在三个视觉演绎推理任务上实现更高的分布外准确率。

AI中文摘要:

视觉语言模型(VLM)已被广泛应用于各种推理任务,但尚不清楚它们在分布偏移下能否稳健推理。本文研究协变量偏移,即感知输入分布变化而底层预测规则不变的情况。为探究此问题,我们考虑视觉演绎推理任务,模型需根据图像和定义在图像中对象概念上的逻辑规则回答查询。实验发现,通过基于梯度的端到端训练微调的VLM能实现高分布内准确率,但在此类偏移下无法泛化,表明微调不能可靠地诱导底层推理函数。这促使我们从神经符号角度将感知与推理分离。然而,我们进一步观察到,近期依赖黑箱组件进行推理的神经符号方法在不同任务间仍可能表现出不一致的稳健性。为解决此问题,我们提出VLC,一种结合基于VLM的概念识别与基于电路的符号推理的神经符号方法。具体而言,任务规则被编译成符号程序(即电路),该程序在VLM识别的对象概念上精确执行规则。在三个具有不同规则集的简单视觉演绎推理任务上的实验表明,VLC在分布外数据上始终比其它推理范式获得更高的任务准确率。代码见此链接。

英文摘要:

Vision-Language Models (VLMs) have been applied to a wide range of reasoning tasks, yet it remains unclear whether they can reason robustly under distribution shifts. In this paper, we study covariate shifts in which the perceptual input distribution changes while the underlying prediction rules do not. To investigate this question, we consider visual deductive reasoning tasks, where a model is required to answer a query given an image and logical rules defined over the object concepts in the image. Empirically, we find that VLMs fine-tuned through gradient-based end-to-end training can achieve high in-distribution accuracy but fail to generalize under such shifts, suggesting that fine-tuning does not reliably induce the underlying reasoning function. This motivates a neuro-symbolic perspective that decouples perception from reasoning. However, we further observe that recent neuro-symbolic approaches that rely on black-box components for reasoning can still exhibit inconsistent robustness across tasks. To address this issue, we propose VLC, a neuro-symbolic method that combines VLM-based concept recognition with circuit-based symbolic reasoning. In particular, task rules are compiled into a symbolic program, specifically a circuit, which executes the rules exactly over the object concepts recognized by the VLM. Experiments on three simple visual deductive reasoning tasks with distinct rule sets show that VLC consistently achieves higher task accuracy on out-of-distribution data than other reasoning paradigms. Code is available at https://github.com/uiuctml/VLC.

补充信息

↑