arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

图表超越像素:探测逐层图表理解与编辑

Charts Are Beyond Pixels: Probing for Layer-Wise Chart Understanding and Editing

Xiaochuan Zhong, Yifan Hou, Chenxi Pang, Shaobo Cui

arXiv 2609.08657首次发表:更新:

发表机构

DeepDelta Lab, School of Artificial Intelligence, Shanghai Jiao Tong University; ETH Zurich; Google DeepMind(上海交通大学人工智能学院DeepDelta实验室; 苏黎世联邦理工学院; 谷歌DeepMind)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有图表基准忽略逐层行为的问题,提出LayerWiseBench基准,基于图层归因、绑定和可见性排序评估模型,发现可见性排序任务最具挑战性,需显式建模组件身份与可见性关系。

AI 中文摘要

图表是结构化的视觉组合,其元素具有不同的功能角色、语义对应关系和可见性关系。这种结构化视角促使我们评估模型是否能在图层级别理解和操作图表。然而,现有的图表基准主要评估最终输出的正确性或保真度,并未直接评估这些逐层行为。我们提出了LayerWiseBench,一个围绕三个核心概念(图层归因、图层绑定和可见性排序)组织的基准,这些概念构成了其图表理解和图表编辑评估的结构。LayerWiseBench由可执行的图表程序生成,将每个渲染的图表与空间对齐的逐层RGBA资产以及用于功能角色、语义绑定和可见性关系的构造派生标签配对。基于这种逐层表示,我们推导出受控的理解问题、编辑目标、参考图像和评估区域。它包含14种图表范式下的2,800个源图表,从中我们推导出7,329个逐层理解问题和53,791个指令引导的编辑变体。在评估的视觉语言模型(VLMs)中,Qwen3.5-27B在问答宏平均上得分最高,在图层归因上达到93.04%的准确率,在图层绑定上达到97.46%,但在可见性排序上仅为61.46%。在四个评估的图像编辑器中,整体mIoU范围从1.49%到4.93%,且对于每个编辑器,可见性约束编辑的mIoU最低,范围从0.37%到2.00%。综合来看,这些结果识别出涉及重叠组件之间前后关系的任务在理解和编辑中反复出现,这促使对组件身份和可见性关系进行更显式的建模。

英文摘要

Charts are structured visual compositions whose elements have distinct functional roles, semantic correspondences, and visibility relations. This structural view motivates evaluating whether models can understand and manipulate charts at the layer level. Existing chart benchmarks, however, primarily assess the correctness or fidelity of final outputs and do not directly evaluate these layer-wise behaviors. We present LayerWiseBench, a benchmark organized around three core concepts, layer attribution, layer binding, and visibility ordering, that structure its chart-understanding and chart-editing evaluations. Generated from executable chart programs, LayerWiseBench pairs each rendered chart with spatially aligned per-layer RGBA assets and construction-derived labels for functional roles, semantic bindings, and visibility relations. From this layer-wise representation, we derive controlled understanding questions, editing targets, reference images, and evaluation regions. It contains 2,800 source charts across 14 chart paradigms, from which we derive 7,329 layer-wise understanding questions and 53,791 instruction-guided editing variants. Among the evaluated VLMs, Qwen3.5-27B, which achieves the highest QA macro-average, obtains 93.04% accuracy on layer attribution and 97.46% on layer binding, but only 61.46% on visibility ordering. Across the four evaluated image editors, overall mIoU ranges from 1.49% to 4.93%, and visibility-constrained edits have the lowest mIoU for every editor, ranging from 0.37% to 2.00%. Taken together, these results identify tasks involving front-to-back relations between overlapping components as a recurring challenge across understanding and editing, motivating more explicit modeling of component identity and visibility relations.

Comments25 pages, 9 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑