arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ChartRevise:通过代码进行精确图表编辑的数据集与评估协议

ChartRevise: A Dataset and Evaluation Protocol for Exact Chart Editing via Code

Jiaxiang Tang, Yi Zhou, Chad DeLuca, Rogerio Feris, Ahmed Khalil Omran, Zhi-Li Zhang, Pengyuan Li, Ali Anwar

arXiv 2609.38642首次发表:更新:

发表机构

University of Minnesota, Twin-Cities; IBM Research; Horizon School of Digital Technologies(明尼苏达大学双城分校; IBM研究院; 地平线数字技术学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有图表编辑基准无法区分请求完成与耦合更新及多余更改的问题,提出ChartRevise数据集与评估协议,基于图形语法构建92,438条记录,覆盖344种编辑类型,通过无参考协议精确评估,微调后需求召回率提升16%,精确编辑率提升22%。

AI 中文摘要

图表编辑需要跨模态的编辑定位,即在绘制图表的代码中实现所请求的视觉变化,同时进行必要的相关更新,且不改变无关内容。现有基准侧重于代码可执行性或图表质量,但其指标未能清晰区分请求完成、遗漏的耦合更新和多余更改。我们引入了ChartRevise,一个结构化的数据集和评估协议,用于精确的基于程序的图表编辑。在数据集构建方面,我们基于图形语法系统地覆盖图表编辑操作,使用源程序检查来验证它们在不同图表类型和库中的适用性。为提高编辑精确性,我们的流程检查各项要求,并在未满足时引导修复或排除。最终数据集包含92,438条记录,覆盖20种图表类型和三个绘图库中的344种编辑类型。在评估方面,我们的无参考协议分别衡量原子需求完成度,识别多余更改,并检测遗漏的耦合更新。这些检查与成功执行和渲染相结合,以确定精确编辑的成功。在五个模型和四个外部基准上,微调在平均需求召回率上相对提升了16%,在平均精确编辑率上相对提升了22%。

英文摘要

Chart editing requires cross-modal edit grounding, realizing a requested visual change in the code that draws it, with necessary related updates and without altering unrelated content. Existing benchmarks emphasize either code executability or chart quality, but their metrics do not clearly distinguish request completion from missed coupled updates and gratuitous changes. We introduce ChartRevise, a structured dataset and evaluation protocol for exact program-grounded chart editing. For dataset construction, we build on the grammar of graphics to systematically cover chart-editing operations, using source-program checks to verify their applicability across chart types and libraries. To improve edit exactness, our pipeline checks individual requirements and guides repair or exclusion when they are unmet. The resulting dataset contains 92,438 records covering 344 edit types across 20 chart types and three plotting libraries. For evaluation, our reference-free protocol separately measures atomic requirement completion, identifies gratuitous changes, and detects missed coupled updates. These checks are combined with successful execution and rendering to determine exact-edit success. Across five models and four external benchmarks, fine-tuning yields relative gains of 16\% in mean requirement recall and 22\% in mean exact-edit rate.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑