VisEditBench:视觉语言模型能否基于多模态反馈编辑可视化代码?
VisEditBench: Can Vision-Language Models Edit Visualization Code from Multimodal Feedback?
浏览论文内容
中文总结 AI 辅助
本研究推出可视化代码编辑基准VisEditBench,评估20种VLMs的编辑能力,提出基于渲染的VisEditAgent框架,显著提升了编辑任务的通过率。
中文摘要 AI 辅助
视觉语言模型(VLMs)已展现出根据文本或视觉规范生成可视化代码的强大能力。然而,现实中的可视化创作本质上是迭代的:用户经常需要修改现有可视化内容,以修复有缺陷的图表或使其适配期望的样式。现有的基准主要评估从零开始的生成任务,而基于多模态反馈的可视化代码编辑任务在很大程度上未被探索。我们推出VisEditBench,这是一个包含1395个人类标注的可视化代码编辑任务的基准,基于真实的可视化工作流和失败案例构建。VisEditBench涵盖两种实际设置:一是反馈引导修复,即模型利用有缺陷或标注过的图表结合文本反馈来修改可视化代码;二是参考引导重设样式,即模型修改代码以匹配目标图表图像。对20种最先进的VLMs的评估显示,可视化代码编辑仍具有挑战性:Claude-4.6-Sonnet的整体通过率达到74.46%,表现最佳,而大多数开源模型的通过率仍低于50%。在基于视觉的样式适配任务上,性能尤为薄弱,其中Claude-4.6-Sonnet仅达到55.71%。为建立强基线,我们进一步提出VisEditAgent,这是一个基于渲染的编辑框架,通过迭代生成、执行、验证和细化候选编辑来完成任务。该框架基于GPT-4o构建,将整体通过率从55.75%提升至67.99%,证明了基于渲染的反馈对于忠实的可视化编辑的重要性。我们将在该httpsURL发布VisEditBench。
英文摘要
Vision-language models (VLMs) have shown strong capabilities in generating visualization code from textual or visual specifications. However, real-world visualization authoring is inherently iterative: users frequently revise existing visualizations to repair flawed charts or adapt them to desired styles. Existing benchmarks primarily evaluate generation from scratch, leaving visualization code editing from multimodal feedback largely unexplored. We introduce VisEditBench, a benchmark of 1,395 human-annotated visualization code-editing tasks grounded in realistic visualization workflows and failure cases. VisEditBench covers two practical settings: feedback-guided repair, where models revise visualization code using buggy or marked charts together with textual feedback, and reference-guided restyling, where models modify code to match a target chart image. Evaluating 20 state-of-the-art VLMs reveals that visualization code editing remains challenging: Claude-4.6-Sonnet achieves the best overall pass rate of 74.46%, while most open-source models remain below 50%. Performance is particularly weak on visually grounded style adaptation, where Claude-4.6-Sonnet achieves only 55.71%. To establish a strong baseline, we further propose VisEditAgent, a render-grounded editing framework that iteratively generates, executes, validates, and refines candidate edits. Built on GPT-4o, VisEditAgent improves overall pass rate from 55.75% to 67.99%, demonstrating the importance of render-grounded feedback for faithful visualization editing. We will release VisEditBench at https://github.com/vis-nlp/VisEditBench.
发表机构
- York University(约克大学)
- Nanyang Technological University(南洋理工大学)
- Salesforce AI Research(Salesforce人工智能研究)
机构由 AI 辅助整理,请以论文原文为准。