WidgetVA:面向智能体可视化分析的小部件中心框架与基准
WidgetVA: A Widget-Centric Framework and Benchmark for Agentic Visual Analytics
查看机构详情
- The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
- Nanyang Technological University(南洋理工大学)
- Sun Yat-sen University(中山大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文提出WidgetVA框架,将可视化分析组件标准化为小部件,并构建WidgetVABench基准,以评估视觉语言模型作为自主VA操作者的能力,实验表明框架有效但存在局限。
中文摘要 AI 辅助
可视化分析(VA)通过交互式可视化实现意义建构,但有效的分析通常需要专家将高层意图转化为长序列的界面操作,并迭代地解读视觉反馈。我们研究现代视觉语言模型(VLM)是否能承担这一角色,成为自主的VA操作者,即观察界面、规划多步探索、执行交互,并根据中间视觉反馈进行调整。为了支持系统性的开发与评估,我们首先引入WidgetVA,一个以小部件为中心的智能体VA框架,它将交互组件标准化为结构化的小部件,并配备统一的动作(如过滤和缩放)和感知查询(如选择摘要)API。这种标准化支持两种系统构建模式:包装现有VA系统使其无需重建即可被智能体操作,以及从小部件作为模块化构建块组合新系统。为了帮助智能体跨小部件协调而非从头规划每次交互,每个小部件还封装了可复用的分析工作流,为智能体提供了比一组可调用函数更丰富的规划依据。基于此框架,我们提出了WidgetVABench,一个包含单小部件和多小部件VA任务的基准,要求智能体执行多步交互以发现证据并产生可验证的结果。每个任务还提供细粒度的参考标注,使WidgetVABench能够分别对答案、参考轨迹相似度和状态进行评分,而不是将智能体性能合并为一个成功分数。在多个VLM上的实验表明,我们的框架为智能体VA提供了有效的支撑,而诊断性度量则揭示了未来工作中持续存在的局限性。WidgetVA框架和WidgetVABench已在此https URL中发布。
英文摘要
Visual analytics (VA) enables sensemaking through interactive visualization, but effective analysis often requires experts to translate high-level intents into long sequences of interface operations and iteratively interpret visual feedback. We study whether modern vision-language models (VLMs) can take on this role as autonomous VA operators that observe the interface, plan multi-step exploration, execute interactions, and adapt based on intermediate visual feedback. To support systematic development and evaluation, we first introduce WidgetVA, a widget-centric agentic VA framework that standardizes interactive components as structured widgets with unified action (e.g., filter and zoom) and perception-query (e.g., selection summaries) APIs. This standardization supports two modes of system construction: wrapping an existing VA system to make it agent-operable without rebuilding it, and composing a new system from widgets as modular building blocks. To help agents coordinate across widgets rather than plan each interaction from scratch, each widget further packages reusable analytical workflows, giving agents more than a bare set of callable functions to plan over. Building on this framework, we present WidgetVABench, a benchmark of single- and multi-widget VA tasks that require agents to perform multi-step interactions to uncover evidence and produce verifiable results. Each task also provides fine-grained reference annotations so that WidgetVABench can score Answer, Reference Trace Similarity, and State separately rather than collapsing agent performance into one success score. Experiments across multiple VLMs show that our framework provides an effective scaffold for agentic VA, while the diagnostic measures expose persistent limitations for future work. The WidgetVA framework and WidgetVABench have been released in https://github.com/Hiverwin/widgetva.