arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13745cs.CL

VLM图表阅读内部机制:跨空间与深度追踪垂直条形图的值读取

Inside VLM Chart Reading: Tracing Value Reading from Vertical Bar Charts Across Space and Depth

  • Research Center for Social Computing and Interactive Robotics(社会计算与交互机器人研究中心)
  • Harbin Institute of Technology, China(哈尔滨工业大学)

机构由 AI 辅助整理,请以论文原文为准。

Tianhao Niu, Qingfu Zhu, Wanxiang Che

中文总结 AI 辅助

本研究通过受控反事实激活修补,在Qwen2.5VL-7B-Instruct和InternVL3.5-8B中追踪垂直条形图值读取,揭示条顶区域、图例与提示系列状态的作用,为定位精确值读取的内部计算提供初步因果证据。

中文摘要 AI 辅助

视觉语言模型(VLM)能够准确回答图表问题,但输出准确性并不能展示它们如何组合恢复精确值所需的证据。我们通过受控的反事实激活修补,在Qwen2.5VL-7B-Instruct和InternVL3.5-8B中研究垂直条形图的值读取。该研究连接了三个分析:(1)单因素结果显示,改变的条顶区域比未改变的条身恢复了更多的答案偏好,尽管前者包含更少的视觉标记。图例和系列相关状态也比条几何和轴刻度状态更早失去局部可恢复性。(2)在交接分析中,恢复从早期层的视觉图例区域转移到中间层的提示系列位置。重置提示系列状态选择性地减少图例来源的救援,支持其作为部分中介的作用。(3)在因子分析中,两个模型都能使用来自不同供体的几何和刻度状态来偏向组合目标。当状态来自不同供体或同一图像时,InternVL表现相似,而Qwen在分离供体时显示出较低的恢复,表明其更高的上下文敏感性。综合来看,这些结果为定位支持精确条值读取的内部计算提供了初步的因果证据。

英文摘要

Vision--language models (VLMs) can answer chart questions accurately, but output accuracy does not show how they combine the evidence needed to recover an exact value. We study vertical-bar value reading with controlled counterfactual activation patching in Qwen2.5VL-7B-Instruct and InternVL3.5-8B. The study connects three analyses: (1) The single-factor results show that the changed bar-top region restores much more answer preference than the unchanged bar body, despite containing fewer visual tokens. Legend- and series-related states also lose local recoverability earlier than bar-geometry and axis-scale states. (2) In the handoff analysis, restoration shifts from visual legend regions in early layers to prompt-series positions in middle layers. Resetting the prompt-series state selectively reduces legend-source rescue, supporting its role as a partial mediator. (3) In the factorial analysis, both models can use geometry and scale states from separate donors to favor the combined target. InternVL performs similarly when the states come from separate donors or one image, while Qwen shows lower restoration for separate donors which suggests higher context sensitivity. Together, these results provide preliminary causal evidence for localizing the internal computations that support exact bar-value reading.

↑