arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从像素到编码:评估多模态大语言模型的图形复现能力

From Pixel to Coding: Evaluating the Figure Reproduction Capabilities of MLLMs

Zijian Chen, Zhengyu Chen, Bohan Liang, Lirong Deng, Yushuo Zheng, Yanwei Jiang, Qi Jia, Kaiwei Zhang, Wenjun Zhang, Guangtao Zhai

arXiv 2610.10066首次发表:更新:

发表机构

Institute of Image Communication and Information Processing, Shanghai Jiao Tong University; Shanghai Artificial Intelligence Laboratory; Macao Polytechnic University(上海交通大学图像通信与信息处理研究所; 上海人工智能实验室; 澳门理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多模态大模型图形复现评估缺失问题,提出FigCodeBench框架,含6194实例、7类功能、4种语言,经24个模型实验发现普遍非线性性能断崖。

AI 中文摘要

多模态大语言模型(MLLMs)在视觉理解和代码生成方面均展现出令人印象深刻的能力。然而,现有基准测试通常将这两种模态孤立评估,缺乏对其统一能力的专门评测,即模型如何感知复杂视觉结构并将其综合为精确、可执行的代码。此外,当前的视觉代码生成基准往往依赖单一编程环境中的简化布局,未能充分评估真正的统一多模态推理。为弥补这一空白,我们提出了FigCodeBench,一个用于严格评估多模态大语言模型在图形复现任务上表现的综合框架,该框架整合了多模态理解与生成。我们首先设计了一个系统化的数据集构建流程,最终生成了总计6,194个实例,涵盖7个功能类别和4种编程语言类型。我们进一步将图形复现划分为三个层级,并进行视觉与代码复杂度建模,特别针对复杂结构推理、不同宽高比和密集几何约束。我们引入了一个多维评估协议,涵盖视觉保真度和语法同构性,该协议与平均机器意见分数(MMOS)和人类偏好高度一致。基于我们的框架,我们对24种广泛使用的专有和开源多模态大语言模型(例如Gemini 3.1 Pro、GPT-5.4和Kimi-K2.5)进行了大量实验,观察到所有模型在不同编程语言和难度场景下均出现普遍的非线性性能断崖,并获得了一些洞见,例如在刚性声明式语言中指标显著下降。

英文摘要

Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in both visual understanding and code generation. However, existing benchmarks typically evaluate these two modalities in isolation, lacking a dedicated assessment of their unification, i.e., how a model can perceive complex visual structures and synthesize them into precise, executable code. Moreover, current visual code generation benchmarks often rely on simplified layouts within single programming environments, falling short of evaluating true unified multimodal reasoning. To bridge this gap, we propose FigCodeBench, a comprehensive framework for rigorously evaluating MLLMs on figure reproduction, integrating multimodal comprehension and generation. We first design a systematic dataset construction pipeline, resulting in a total of 6,194 instances that cover 7 functional categories and 4 types of programming languages. We further categorize figure reproduction into three tiers with visual and code complexity modeling, specifically targeting complex structural reasoning, varying aspect ratios, and dense geometric constraints. We introduce a multi-dimensional evaluation protocol, encompassing visual fidelity and syntactic isomorphism, that aligns highly with the Mean Machine Opinion Score (MMOS) and human preferences. Based on our framework, we conducted extensive experiments on 24 widely used proprietary and open-source MLLMs (e.g., Gemini 3.1 Pro, GPT-5.4, and Kimi-K2.5), where we observed a universal, non-linear performance cliff across different programming languages and difficulty scenarios for all models, and gained several insights, such as the significant metric decline in rigid declarative languages.

Comments46 pages, 18 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑