发表机构
Pandita AI Inc.(潘迪塔人工智能公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究推出首个专门评估LLMs数学图表生成能力的基准Math-Vision Diagrams,涵盖文本到代码与图像范式,测试发现LLMs在该任务上存在困难,相关资源将开源。
AI 中文摘要
从文本提示生成数学上精确的图表,已成为大型语言模型(LLMs)一项关键但未被充分探索的能力。这一能力受到课程准备、习题集自动排序和科学出版领域研究人员的关注。LLMs要实现这一点,需要空间推理、数学推理和渲染系统之间的完美协调。现有基准如MathVision、MathVista是为数学推理设计的,DiagramGenBenchmark、MermaidSeqBench则针对通用图表生成,尚无研究提供标准化的提示-图像对,专门用于评估LLMs的数学图表生成能力,涵盖文本到代码和文本到图像两种范式。我们推出Math-Vision Diagrams,首个专门设计用于评估LLMs数学图表生成能力的基准,也是首个在统一设置中同时评估文本到代码和文本到图像生成范式的基准,不依赖底层编程语言或模型类型。基于Math-Vision基准,我们从3040张高质量竞赛题图像中选出2920张包含必要视觉上下文的子集。我们提出一种结合LLMs集成与主题专家(SME)审核的新型流程,以及一套评估指标。通过在该基准上测试多个领先模型,我们证明LLMs在数学图表生成方面存在困难。所有代码、数据、审核流程和评估脚本将完全开源。
英文摘要
The generation of mathematically precise diagrams from tex- tual prompts has emerged as a critical yet underexplored capability of Large Language Models (LLMs). This has been of interest to researchers in the areas of curriculum preparation, automated ranking of problem sets, and scientific publishing. For LLMs to achieve this, it requires per- fect coordination between Spatial Reasoning, Mathematical Reasoning, and Rendering systems. While existing benchmarks such as MathVision, MathVista are built for Math Reasoning or DiagramGenBenchmark, Mer- maidSeqBench on general purpose diagram generation, no prior work provides a standardized set of prompt, image pairs that can be used to evaluate the LLMs specifically on math diagram generation. This includes fields that span both both text-to-code and text-to-image paradigms. We introduce Math-Vision Diagrams, the first benchmark specifically designed to evaluate LLMs on mathematical diagram generation, and the first to assess text-to-code and text-to-image generation paradigms together in a single unified setting, agnostic of the underlying coding lan- guage or model type. Building on the Math-Vision benchmark, we select a subset of 2920 images out of 3040 from high-quality competition problems with essential visual context. A novel pipeline combining an ensemble of LLMs with Subject Matter Expert (SME) curation is presented, together with a suite of evaluation metrics. Testing several leading models against this benchmark, we demonstrate that LLMs struggle with math diagram generation. All code, data, curation pipeline, and evaluation scripts will be fully open-sourced.
CommentsAccepted at ICANN 2026