arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Edit2TikZ:面向基于TikZ的科学图像编辑的全面且具有挑战性的基准

Edit2TikZ: A Comprehensive and Challenging Benchmark for Scientific Figure Editing with TikZ

Zongyun Zhang, Jiacheng Ruan, Xian Gao, Ruizhu Zhou, Lingcheng Meng, Lining Hu, Ting Liu, Yuzhuo Fu

arXiv 2608.13441首次发表:更新:

AI 中文总结

本文推出科学图像编辑基准Edit2TikZ,评估主流多模态大语言模型发现其性能不足,通过构建混合训练集并采用课程学习,可显著提升紧凑模型的编译成功率。

AI 中文摘要

尽管多模态大语言模型(MLLMs)在视觉理解和图形代码生成方面展现出巨大潜力,但通过代码编辑科学图像面临更大挑战:模型必须同时恢复视觉结构、定位请求的修改、生成可编译代码,并保留所有无关内容。现有TikZ基准主要聚焦于图像重建和生成,却很少系统评估基于指令的科学图像编辑及可编译代码。我们推出Edit2TikZ,这是一个面向科学图像编辑任务的全面基准,包含1548个多样化且高质量样本。Edit2TikZ结合了真实场景和受控合成编辑案例,支持文本和视觉定位请求,还包含多步骤编辑,每个步骤都有步骤级注释。我们进一步构建了人类对齐的评估框架,用于衡量请求的编辑是否完成且无关内容是否保留。利用Edit2TikZ,我们评估了14个主流MLLMs,发现当前系统仍不可靠:平均而言,专有模型的编译成功率仅为75%,在图像恢复和编辑正确性方面仍存在局限,而规模低于9B的紧凑模型在遵循指令和完成图像生成方面表现更差。因此,我们构建了混合训练集TikZEditMix,并对紧凑模型采用“先重建后编辑”的课程学习方法。在Qwen3.5-4B上,该训练方法将编译成功率从45.35%提升至83.40%,在我们提出的评估指标上平均提升了18.7个百分点。代码和数据将发布在该https URL。

英文摘要

Although multimodal large language models (MLLMs) have shown substantial potential in visual understanding and graphic code generation, editing scientific figures through code presents a greater challenge: a model must jointly recover visual structure, ground the requested change, generate compilable code, and preserve all unrelated content. While existing TikZ benchmarks mainly focus on figure reconstruction and generation, few systematically evaluate instruction-guided scientific figure editing with compilable code. We introduce Edit2TikZ, a comprehensive benchmark for scientific figure editing tasks, featuring 1,548 diverse and high-quality samples. Edit2TikZ combines real-world and controlled synthetic edit cases, supports both textual and visual localization request, and contains multi-step editing, each with step-level annotations. We further construct a human-aligned evaluation framework to measure whether a requested edit is completed while irrelevant content is preserved. Utilizing Edit2TikZ, we evaluate 14 mainstream MLLMs and find that current systems remain unreliable: on average, proprietary models achieve a compilation success rate of merely 75% and remain limited in both figure restoration and edit correctness, while compact models below 9B struggle further with instruction following and complete figure generation. Therefore, we build a mixed training set TikZEditMix and adopt reconstruction-then-editing curriculum learning for compact models. On Qwen3.5-4B, this training improves the compilation success rate from 45.35% to 83.40% and yields an average improvement of 18.7 points across our proposed evaluation metrics. The code and data will be released at https://github.com/Solunny/Edit2TikZ.

Comments9 pages, 6 figures, work in progress

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑