arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13267cs.CV

科学图表理解的结构-标记证据锚定推理

Structure-Token Evidence-Anchored Reasoning for Scientific Chart Understanding

  • University of Brasilia(巴西利亚大学)

机构由 AI 辅助整理,请以论文原文为准。

Alberlucia Rafael Soarez, Camila Ferreira, Daniel Kim, Mariana Costa, Alejandro Torres

中文总结 AI 辅助

针对科学图表理解,提出STEER方法,通过结构图编码、证据锚定推理和弱解析器对齐,在多个基准上超越现有模型,显著提升推理准确率。

中文摘要 AI 辅助

科学图表通过坐标轴、图例和几何标记编码数量信息,然而大型视觉语言模型仍将它们视为自然照片。视觉上下文示例无法揭示坐标框架;不受约束的思维链可能说出一个从未从条形图中读取的看似合理的数字。我们提出STEER(结构-标记证据锚定推理),该方法冻结Llama-3.2-Vision编码器并插入三个模块:图表结构图编码器(CSGE)用于绑定刻度、图例项和标记;证据锚定步骤推理(EASR)强制每个算术步骤引用图节点;弱解析器强推理器对齐(WPSR)仅将专用表格提取器用作节点属性的教师。在ChartQA上,STEER达到82.70的平均宽松准确率,而ChartGemma为80.16,在同一混合数据上训练的LLaVA-CoT骨干为76.40。在CharXiv推理(33.60对比InternVL Chat V1.5的29.20)和ChartQAPro CoT(40.70对比Qwen2-VL-7B的37.17)上,增益扩大,OCR捷径消失。消融实验表明,丢弃节点序列化或数值候选约束会抵消大部分推理提升。

英文摘要

Scientific charts encode quantities in axes, legends, and geometric marks, yet large vision-language models still treat them as natural photographs. Visual in-context examples do not expose the coordinate frame; unconstrained chain-of-thought can name a plausible number that was never read from a bar. We present STEER (Structure-Token Evidence-anchored Reasoning), which freezes a Llama-3.2-Vision encoder and inserts three modules: a chart structure graph encoder (CSGE) that binds ticks, legend items, and marks; evidence-anchored step reasoning (EASR) that forces every arithmetic step to cite a graph node; and weak-parser strong-reasoner alignment (WPSR) that uses a specialized table extractor only as a teacher of node attributes. On ChartQA, STEER reaches 82.70 average relaxed accuracy versus 80.16 for ChartGemma and 76.40 for a LLaVA-CoT backbone trained on the same mix. Gains widen on CharXiv reasoning (33.60 vs. 29.20 InternVL Chat V1.5) and ChartQAPro CoT (40.70 vs. 37.17 Qwen2-VL-7B), where OCR shortcuts disappear. Ablations show that dropping node serialization or numeric candidate constraints undoes most of the reasoning lift.

↑