arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.00948cs.CVcs.AIcs.CL

从术语到图表:面向科学图表理解的视觉指令生成

From Terminology to Diagrams: Visual-Instruction Generation for Scientific Diagram Understanding

发表机构语言技术研究实验室
查看机构详情
  • Language Technology Research Laboratory(语言技术研究实验室)

机构由 AI 辅助整理,请以论文原文为准。

Raul Ortega, José Manuel Gómez-Pérez

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出基于科学课程术语的视觉指令生成框架,构建SciGram数据集,微调模型后在图表问答基准上取得性能提升,还增强LLaVA OneVision达到新SOTA,将发布数据集与模型。

中文摘要 AI 辅助

视觉语言模型(VLMs)在自然图像的视觉问答中已展现出强大性能,但在科学图表上仍存在困难——科学图表旨在传达功能或关系意义,而非字面场景。为此,我们提出一种利用科学课程术语生成大规模图表关联指令数据的框架。该方法系统提取领域概念、合成原子事实、从网络检索相关图表,并以图表标题和多项选择题形式生成多模态监督。通过该流程,我们构建了SciGram数据集,涵盖生命科学、地球科学和物理科学领域,包含超19.4万张图表及140万条视觉指令。尽管依赖有噪声的网络数据和合成标注,在SciGram上微调的模型在图表基准测试(包括TQA、ScienceQA和AI2D)中仍取得显著提升,使用更少训练实例即可超越或匹配最先进的VLMs;此外,用SciGram增强现有模型(如LLaVA OneVision)在图表问答上达到新的最先进性能。我们的结果表明,基于术语的指令生成是提升科学领域视觉语言推理的有效通用策略,为支持未来科学图表理解研究,我们发布SciGram数据集及模型。

英文摘要

Vision-language models (VLMs) have demonstrated strong performance in visual question answering with natural images. However, they continue to struggle with scientific diagrams, which are designed to convey functional or relational meaning rather than literal scenes. We therefore introduce a framework for generating large-scale diagram-grounded instruction data by leveraging terminology derived from scientific curricula. Our approach systematically extracts domain concepts, synthesizes atomic facts, retrieves relevant diagrams from the web, and generates multimodal supervision in the form of diagram captions and multiple-choice questions. Using this pipeline, we construct SciGram, a dataset of over 194K diagrams and 1.4M visual instructions across life, earth, and physical sciences. Despite relying on noisy web data and synthetic annotations, models fine-tuned on SciGram achieve substantial improvements on diagram-centric benchmarks, including TQA, ScienceQA, and AI2D, outperforming or matching state-of-the-art VLMs while using fewer training instances. Furthermore, augmenting existing models such as LLaVA OneVision with SciGram establishes new state-of-the-art performance on diagram question answering. Our results highlight the effectiveness of terminology-grounded instruction generation as a general strategy for improving vision-language reasoning in scientific domains. To support future research in scientific diagram understanding, we release both the SciGram dataset and models.

补充信息

↑