arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.10301cs.CVcs.CL

ChartSync:视觉逻辑级联图表编辑的基准测试

ChartSync: A Benchmark for Visuo-Logical Cascading Chart Editing

  • Jianghan University(江汉大学)
  • HiThink Research(海思睿研)
  • University of Science and Technology of China(中国科学技术大学)
  • Southeast University(东南大学)
  • Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

Jiakang Yu, Yixuan Chai, Tianci Wang, Rihui Jin, Guangkai Xu, Hongtao Deng, Xun Zhu, Wang Gao, Xinrun Guo, Haipang Wu

中文总结 AI 辅助

研究针对生成式图像编辑模型处理结构化统计图的困难,提出视觉逻辑级联编辑任务。引入ChartSync基准测试,含多种图表类别和任务类型。通过两层框架评估模型,发现开源与专有模型能力差距,分析误差以指导多模态架构发展,数据集和代码已公开。

中文摘要 AI 辅助

生成式图像编辑模型在处理需要几何同步的数据修改的结构化统计图时存在困难。我们将此任务形式化为视觉逻辑级联编辑(VLCE)。现有方法局限于局部文本替换,难以处理依赖感知级联更新。为系统评估此能力,我们引入ChartSync,通过程序化渲染管道构建的专家验证基准,保证与地面真值的确定性视觉逻辑耦合。ChartSync包含9个图表类别和4种任务类型的870个三元组,包括235个几何耦合VLCE实例。我们通过结合客观视觉指标与视觉语言模型判断范式的两层框架评估这些实例。评估14个图像编辑模型和一个代码介导管道发现了细微的能力差距:大多数开源模型在几何同步方面严重下降,只有两个前沿专有模型显示出新兴的VLCE能力,其残余误差主要涉及语义隔离和背景损坏。我们的详细误差分析解构了这些失败范式,以识别指导未来多模态架构的核心元能力。ChartSync数据集和代码在该https URL上公开发布。

英文摘要

Generative image editing models struggle with structured statistical charts when data modifications require geometric synchronization. We formalize this task as Visuo-Logical Cascading Editing (VLCE). However, existing methods remain confined to localized text substitutions and struggle with dependency-aware cascading updates. To systematically evaluate this capability, we introduce ChartSync, an expert-validated benchmark constructed via a programmatic rendering pipeline that guarantees deterministic visuo-logical coupling for the ground truth. ChartSync comprises 870 triplets across 9 chart categories and 4 task types, including 235 geometry-coupled VLCE instances that specifically test cascading text-to-geometry synchronization. We further evaluate these instances via a two-tier framework combining objective visual metrics with a vision-language model judge paradigm to assess low-level fidelity alongside multimodal comprehension and reasoning. Evaluating 14 image editing models and one code-mediated pipeline reveals a nuanced capability gap: most open-source models suffer severe drops in geometric synchronization, while only two frontier proprietary models show emerging VLCE capability, with their residual errors mainly involving semantic isolation and background corruption. Our detailed error analysis deconstructs these failure paradigms to identify core meta-abilities for guiding future multimodal architectures. The ChartSync dataset and code are publicly released at https://github.com/kaka-yjk/ChartSyncCodebase.

↑