IChart2Code:面向交互式图表代码生成的多模态大语言模型基准测试
IChart2Code: Benchmarking Multimodal Large Language Models for Interactive Chart Code Generation
浏览论文内容
中文总结 AI 辅助
提出IChart2Code基准,含377个交互式图表代码生成任务及浏览器评估协议,并设计TRAIL双智能体框架,通过轨迹引导修复,在四个MLLM上平均提升各维度3.45-7.89个百分点。
中文摘要 AI 辅助
交互式图表代码生成要求模型复现参考图表的外观和底层数据,并正确实现由指定用户交互触发的状态变化。现有的图表到代码基准侧重于静态输出,缺乏针对交互规范、浏览器执行和交互后验证的任务表示或评估协议。我们提出了IChart2Code,一个包含20种图表形式和13个数据家族中377个任务、涵盖6个家族中1209个交互需求的基准。每个任务提供参考截图、任务局部数据和自然语言交互需求,并以可执行的HTML/JavaScript代码作为目标输出。我们进一步开发了基于浏览器的评估协议,包含可执行性门控和三个基于评分标准的维度:数据保真度、静态视觉正确性和交互正确性。该协议在沙盒浏览器中测试运行时可行性、与任务局部数据的一致性、初始渲染对参考截图的保真度以及交互引起的状态变化。一个基于评分标准的MLLM评判器使用收集的浏览器观察结果评估三个评分维度的任务特定项目,并与裁定的人类标签相比,实现了0.8844的总体项目级F1分数。我们还提出了TRAIL,一个用于交互式图表代码生成的轨迹引导双智能体框架。一个检查器(Inspector)推导任务特定的检查项,在浏览器中执行它们,并使用生成的轨迹来诊断失败并产生结构化的修复反馈。在四个MLLM上平均,TRAIL在四个评估维度上相比直接提示分别提高了7.89、4.98、3.45和4.47个百分点。
英文摘要
Interactive chart code generation requires models to reproduce a reference chart's appearance and underlying data and correctly implement the state changes triggered by specified user interactions. Existing chart-to-code benchmarks focus on static outputs and lack task representations or evaluation protocols for interaction specification, browser execution, and post-interaction verification. We introduce IChart2Code, a benchmark comprising 377 tasks across 20 chart forms and 13 data families, with 1209 interaction requirements in six families. Each task provides a reference screenshot, task-local data, and natural-language interaction requirements, with executable HTML/JavaScript code as the target output. We further develop a browser-based evaluation protocol with an Executability gate and three rubric-guided dimensions: Data Fidelity, Static Visual Correctness, and Interaction Correctness. The protocol tests runtime viability, consistency with task-local data, fidelity of the initial rendering to the reference screenshot, and interaction-induced state changes in a sandboxed browser. A rubric-guided MLLM judge evaluates task-specific items for the three scored dimensions using the collected browser observations and achieves an overall item-level F1 score of 0.8844 against adjudicated human labels. We also propose TRAIL, a trajectory-guided dual-agent framework for interactive chart code generation. An Inspector derives task-specific inspection checks, executes them in the browser, and uses the resulting trajectories to diagnose failures and produce structured repair feedback. Averaged across four MLLMs, TRAIL improves the four evaluation dimensions over direct prompting by 7.89, 4.98, 3.45, and 4.47 percentage points, respectively.
发表机构
- School of Big Data and Software Engineering, Chongqing University(重庆大学大数据与软件工程学院)
机构由 AI 辅助整理,请以论文原文为准。