发表机构
York University(约克大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对语言模型和视觉语言模型生成交互式数据可视化界面的能力缺乏评估的问题,提出VIS-GEN基准,含3042个样本,测试14个模型,发现性能差距大,并设计多阶段生成框架,将最佳模型通过率提升15.9个百分点。
AI 中文摘要
数据可视化是分析推理的核心,但现实世界中的分析越来越需要语言驱动的交互式界面,而非静态图表。尽管近期的大型语言模型和视觉语言模型(LLMs/VLMs)在从自然语言生成静态图表方面展现出潜力,但由于缺乏基准测试,它们生成交互式数据可视化界面的能力在很大程度上仍未得到探索。我们引入了VIS-GEN,一个用于评估LLMs/VLMs如何从自然语言查询生成交互式可视化界面的基准。VIS-GEN包含3,042个样本,涵盖多样化的分析意图,包括数据过滤、时间分析和可视化编辑,每个样本都配有数据集元数据和自然语言查询,旨在反映现实的、目标驱动的数据探索场景。我们对14个最先进的开源和闭源LLMs/VLMs进行了基准测试,揭示了在涉及隐含意图、多种交互替代方案和复杂编辑操作的查询上存在巨大的性能差距和频繁失败,凸显了交互式界面生成是超越静态图表合成的关键开放挑战。为解决这一问题,我们提出了一个结构化的多阶段界面生成框架,将任务分解为可视化设计表示、多个界面候选的生成、约束感知的批判和自我完善。该方法将最佳模型的通过率提高了15.9个百分点,展示了通往更可靠的语言驱动交互式可视化系统的实用路径。我们在https://this URL发布了VIS-GEN。
英文摘要
Data visualization is central to analytical reasoning, but real-world analysis increasingly requires language-driven interactive interfaces rather than static charts. Although recent large language and vision language models (LLMs/VLMs) have shown promise in generating static charts from natural language, their ability to generate interactive data visualization interfaces remains largely unexplored due to the lack of benchmarks. We introduce VIS-GEN, a benchmark for evaluating how well LLMs/VLMs can generate interactive visualization interfaces from natural language queries. VIS-GEN comprises 3,042 samples covering diverse analytical intents, including data filtering, temporal analysis, and visualization editing, each paired with dataset metadata and natural language queries that are designed to reflect realistic, goal driven data exploration scenarios. We benchmark 14 state-of-the-art open-source and closed-source LLMs/VLMs, revealing large performance gaps and frequent failures on queries involving implicit intent, multiple interaction alternatives, and complex editing operations, highlighting interactive interface generation as a key open challenge beyond static chart synthesis. To address this, we propose a structured multi stage interface generation framework that decomposes the task into visualization design representation, generation of multiple interface candidates, constraint-aware critique, and self-refinement. This approach improves the best models pass rate by 15.9 percentage points, demonstrating a practical path toward more reliable language-driven interactive visualization systems. We release VIS-GEN at https://github.com/vis-nlp/VIS-GEN.