arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MV-Bench:用于协同多视图界面构建的多模态大语言模型基准测试

MV-Bench: Benchmarking Multimodal Large Language Models for Coordinated Multi-View Interface Construction

Yue Zhao, Hongxu Liu, Feiyu Wang, Xiaoyu Yang, Tong Ge, Zhen Yang, Chao Wang, Qiong Zeng

arXiv 2607.19910首次发表:更新:

发表机构

Shandong Second Medical University; Shandong University; Bairong Inc.; Ke Holdings Inc.; School of Computer Science and Technology, Shandong University; Shandong Provincial Key Laboratory of Computing-Network Integration, Shandong University(山东第二医科大学; 山东大学; 百融云创科技股份有限公司; 贝壳控股有限公司; 山东大学计算机科学与技术学院; 山东大学计算网络融合技术山东省重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对多模态大语言模型在协同多视图界面构建评估不足的问题,引入MV-Bench基准,通过多阶段管道转换规范为界面,评估五个模型,发现当前MLLMs虽能再现视觉外观,但在数据语义和交互逻辑生成上有限,迭代改进效果不佳。

AI 中文摘要

多模态大语言模型(MLLMs)有望通过直接从视觉设计生成代码来实现可视化开发自动化。然而,现有评估主要集中在单图表生成,忽视了协同多视图界面构建,该领域缺乏专门基准测试。我们引入MV-Bench,用Tableau工作簿文件作为基准来评估MLLMs在协同多视图界面构建方面的能力。开发多阶段管道将规范转换为可执行网络界面。基准包含92个基本界面和1048个实例。评估五个先进MLLMs,结果显示当前MLLMs能再现视觉外观,但在生成数据语义和交互逻辑方面有限,迭代改进可提高代码可执行性,但缩小数据绑定和交互生成差距效果不明显。

英文摘要

Multimodal large language models (MLLMs) are increasingly expected to automate visualization development by generating code directly from visual designs. However, existing evaluations mainly focus on single-chart generation and overlook coordinated multi-view interface construction, which requires joint reasoning about data semantics, view coordination, and interaction logic. Consequently, MLLM capabilities in this setting remain underexplored, and the field lacks a dedicated benchmark for systematic assessment. We introduce MV-Bench, a benchmark for evaluating MLLMs on coordinated multi-view interface construction. Instead of relying on incomplete or inconsistent open-source implementations, we use Tableau workbook files as ground truth because they explicitly encode data bindings, visual mappings, and interactions. We develop a multi-stage pipeline that converts these specifications into executable web interfaces through structured intermediate representations. The benchmark contains 92 base interfaces and 1,048 verified instances created by recombining chart types, datasets, and interaction patterns. Each instance includes executable code, a rendered interface, a dataset, and interaction annotations. We evaluate five state-of-the-art MLLMs in a single-pass setting using metrics for visual fidelity, data binding correctness, and interaction completeness. The strongest model achieves 75.45 percent accuracy in visual layout reproduction, but only 21.71 percent in data binding and 11.68 percent in interaction completeness. These results show that current MLLMs can reproduce visual appearance but remain limited in generating the data semantics and interactive logic required by coordinated multi-view interfaces. Iterative refinement improves code executability but does not substantially reduce the gap in data binding and interaction generation.

CommentsSubmitted to IEEE VIS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑