arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Diagram-MMU:面向科学图表的多模态基准测试

Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams

Weihao Bo, Shan Zhang, Yanpeng Sun, Jie Liu, Yongke Yao, Jinhao Du, Wei He, Kai Zou, Zechao Li, Jingdong Wang

arXiv 2608.12262首次发表:更新:

发表机构

Nanjing University of Science and Technology; Baidu Inc; AIML, Adelaide University; SUTD; Southeast University; East China Normal University; University of Oxford(南京理工大学; 百度公司; 阿德莱德大学AIML; 新加坡科技设计大学; 东南大学; 华东师范大学; 牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文构建了多模态基准测试Diagram-MMU,评估MLLMs的科学图表解析与理解能力,发现图表转代码任务更具挑战,Claude-4.6 Opus在智能体场景下表现最优。

AI 中文摘要

多模态大语言模型(MLLMs)在科学写作与协作方面的能力正在不断提升。例如,OpenAI Prism是一款用于科学写作与协作的免费工作空间,其重要功能之一是将科学图表直接转换为LaTeX TikZ代码。本文构建了名为Diagram-MMU的多模态基准测试,旨在评估MLLMs对科学图表的解析与理解能力。该基准测试包含6个领域的3700份精心整理的图表及18300个人工验证的问题,针对vibe写作工作空间中常见的三项任务对MLLMs进行评估:图表转代码解析、图表转代码编辑以及图表问答,每项任务均设置了智能体(agentic)场景。对12款MLLMs的评估显示,图表转代码任务比图表问答更具挑战性:模型能够对图表进行良好推理,但在解析和编辑方面存在困难,凸显了提升MLLMs图表转代码生成能力的方法的必要性。在智能体场景下,多数模型的解析与编辑性能有所提升,但问答性能出现下降,而Claude-4.6 Opus在三项任务中均实现了持续提升。项目页面:this https URL。

英文摘要

Multimodal Large Language Models (MLLMs) have been growing the capability for scientific writing and collaboration. For example, OpenAI Prism is a free workspace for scientific writing and collaboration. One important feature in Prism is turning scientific diagrams directly into LaTeX TikZ code. In this paper, we build a benchmark, Diagram-MMU, a multi-modal benchmark designed to assess MLLMs' ability for scientific diagram parsing and understanding. Diagram-MMU features 3.7k curated diagrams and 18.3k human-validated questions across six domains. It evaluates MLLMs on three tasks common in vibe writing workspaces: diagram-to-code parsing, diagram-to-code editing, and diagram question answering, alongside agentic settings per task. The evaluation of 12 MLLMs reveals that diagram-to-code tasks are more challenging than diagram question answering: models can reason well over diagrams but struggle to parse and edit them, underscoring the need for methods to enhance MLLMs' capability in diagram-to-code generation. Under agentic settings, most models improve parsing and editing performance but degrade on question answering, while Claude-4.6 Opus consistently improves across all three tasks. Project Page: https://vi-ocean.github.io/projects/diagram-mmu.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑