发表机构
University of Technology Nuremberg; University of Sheffield(纽伦堡工业大学; 谢菲尔德大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对科学图表编辑任务,构建了首个基于修订轨迹的大规模数据集DaEdiTikZ及人工基准,训练了Qwen3.5系列的EdiTikZ模型,其性能在自动与人工评估中优于多数基线,与顶尖模型相当。
AI 中文摘要
视觉语言模型(VLM)在根据文本或图像生成科学图表方面已展现出强大性能。然而,生成可发表的图表需要迭代优化,这使得科学图表编辑成为一项重要却基本未被探索的任务。现有方法依赖成本高昂的专有智能体系统,主要聚焦于评估,或从合成生成的编辑中构建训练监督。相反,我们利用自然存在的科学修订与开发轨迹作为可扩展的监督来源。为此,我们推出DaEdiTikZ——首个基于修订的大规模科学图表编辑数据集,通过从arXiv、GitHub和TeX SE中挖掘39.1万对合理的TikZ编辑对,并基于渲染图表与TikZ代码,借助VLM推断出78.1万条定向编辑指令而构建。我们还推出DaEdiTikZ-Bench,这是一个经人工优化的基准,包含790个实例;并训练了两个基于Qwen3.5的轻量EdiTikZ模型(4B和9B),通过联合学习重建与编辑,再辅以针对渲染保真度与编辑应用的互补奖励的强化学习(RL)进行训练。自动评估显示,我们的9B模型优于所有测试的基线模型;而由9名标注员开展、共4320条评分的人工评估显示,其性能优于GPT-5.6-Sol,与Gemini-3.1-Pro相当。在严重的分布外偏移场景下,其性能仍与GPT-5.6-Sol接近其2000训练序列长度 regime 时的表现相当。模型与数据集将被发布。
英文摘要
Vision-language models (VLMs) have shown strong performance in generating scientific figures from text or images. However, publication-ready figures often require iterative refinement, making scientific figure editing an important yet largely unexplored step toward interactive figure creation. Existing approaches rely on costly proprietary agentic systems, focus primarily on evaluation, or construct training supervision from synthetically generated edits. Instead, we leverage naturally occurring scientific revision and development trajectories as a scalable source of supervision. To this end, we introduce DaEdiTikZ, the first large-scale dataset of revision-derived scientific figure edits, constructed by mining 391K plausible TikZ edit pairs from arXiv, GitHub, and TeX SE and inferring 781K directed edit instructions with a VLM conditioned on rendered figures and TikZ code. We further introduce DaEdiTikZ-Bench, a human-refined benchmark with 690 instances, and train two compact Qwen3.5-based EdiTikZ models (4B and 9B) by jointly learning image-to-TikZ reconstruction and instruction-conditioned editing, followed by reinforcement learning (RL) with complementary rewards for rendered fidelity and edit application. Automatic evaluation places our 9B model above all tested baselines, while human evaluation with 9 annotators and 4,320 ratings places it above GPT-5.6-Sol and on par with Gemini-3.1-Pro. Under severe out-of-distribution shifts, it remains competitive with GPT-5.6-Sol near its 2K training sequence-length regime.
Comments35 pages, 21 figures, and 19 tables. Models and datasets: https://huggingface.co/collections/nllg/editikz . Code: https://github.com/NL2G/EdiTikZ