ChartProbe:通过感知、定位与简单推理开展视觉推理的诊断研究
ChartProbe: A Diagnostic Study on Visual Reasoning through Perception, Grounding, and Simple Reasoning
浏览论文内容
中文总结 AI 辅助
本文提出诊断框架ChartProbe,发现视觉-语言模型可通过单独监督感知、定位等简单技能,在无需复杂推理数据的情况下提升复杂图表推理能力。
中文摘要 AI 辅助
视觉-语言模型(VLMs)在需要对视觉量进行推理的图表问题上仍不可靠,这一缺陷通常被归因于推理能力不足,并通过增加推理监督来解决。本文探究该难点是否源于推理本身,还是推理所依赖的更基础技能:读取绘制元素(感知)、定位元素并将其与标签绑定(定位),以及执行排序、求和、求差等单步计算(简单推理)。本文提出诊断框架ChartProbe,其探针直接由绘制每个图表的代码生成,因此每个标准答案在构造上是精确的,无需人工标注,且能将每个错误归因于单一技能。ChartProbe支持现有工作未尝试的干预方式:不合成复杂推理数据,完全不使用复杂问题和推理轨迹,而是每次针对一种简单技能进行微调,并测量其对未见过的复杂推理问题的迁移效果。在三个开放权重VLMs上,仅对简单技能进行监督,就能在模型从未训练过的复杂推理问题上取得显著提升:当这些技能薄弱且模型能够读取图像时,训练它们无需复杂推理数据成本即可恢复大量复杂推理能力。该提升在三个分布外场景中均成立:未见过的图表类型(饼图)、与本文图像和模板不重叠的人工编写基准(ChartQA)、非图表视觉领域(CLEVR)。因此,复杂视觉推理无需复杂推理监督即可得到提升。
英文摘要
Vision-language models (VLMs) remain unreliable on chart questions that require reasoning over visual quantities, and this weakness is usually attributed to a reasoning deficit and addressed with more reasoning supervision. We ask whether the difficulty lies in reasoning itself, or in the simpler skills that reasoning operates on: reading the plotted elements (\emph{perception}), locating them and binding them to their labels (\emph{grounding}), and performing single-step computations such as ranking, totals, and differences (\emph{simple reasoning}). We introduce \textbf{ChartProbe}, a diagnostic framework whose probes are generated directly from the code that renders each chart, so every gold answer is exact by construction, needs no human annotation, and attributes each failure to a single skill. ChartProbe enables an intervention prior work does not attempt: instead of synthesizing complex-reasoning data, we withhold complex questions and reasoning traces entirely, fine-tune on one simple skill at a time, and measure transfer to held-out complex-reasoning questions. Across three open-weight VLMs, supervising the simpler skills alone produces large gains on complex-reasoning questions the model never trained on: where these skills are weak and the model can be taught to read the image, training them recovers much of complex reasoning at no reasoning-data cost. The gains hold across three out-of-distribution settings: an unseen chart type (pie charts), a human-written benchmark disjoint from our images and templates (ChartQA), and a non-chart visual domain (CLEVR). Complex visual reasoning can therefore improve without complex-reasoning supervision.