arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DataVista:面向数据视频理解的多模态大语言模型诊断工具

DataVista: Diagnosing Multimodal LLMs on Data Video Understanding

Yupeng Xie, Zhenyang Wang, Jiayi Zhu, Yinghao Tang, Zhouan Shen, Yiyu Chen, Yuyu Luo

arXiv 2610.11993首次发表:更新:

发表机构

The Hong Kong University of Science and Technology (Guangzhou); Zhejiang University(香港科技大学(广州); 浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究人员构建首个数据视频理解基准DataVista,评估19种主流多模态大语言模型,发现最优模型Gemini-3.1-Pro准确率70.0%,远低于人类,揭示模型在因果推理等任务的不足及帧数、字幕的影响。

AI 中文摘要

数据视频是一种将数据可视化与视频叙事相结合的媒体形式,广泛应用于新闻报道和商业分析。与通用视频理解相比,数据视频理解更强调准确读取动画图表中的数据、整合不同图表及不同时间节点的证据,以及理解叙事组织和视觉设计如何传递信息。然而,现有的基准测试要么针对通用视频,要么针对静态图表,数据视频理解尚未得到系统性评估。我们提出DataVista,这是首个面向数据视频理解的基准测试,包含961个真实世界数据视频和6775个评估问题,这些问题基于三级递进能力框架(数据感知、时序推理、叙事理解)组织,涵盖五个主题领域的10种细粒度问题类型。对19种主流多模态大语言模型(MLLM)的系统性评估显示,表现最佳的模型Gemini-3.1-Pro的总体准确率为70.0%,仍远低于人类专家的表现,其中模型在因果推理和叙事结构任务上表现最差。增加帧数和添加字幕主要有益于数据感知和时序推理,对叙事理解的提升有限。对模型响应的进一步分析还识别出图表读取、证据判断和指令理解方面的典型失败模式。该基准测试可通过此https URL获取。

英文摘要

Data video is a media form that integrates data visualization with video narrative, widely adopted in news reporting and business analysis. Compared with general video understanding, data video understanding places greater emphasis on accurately reading data from animated charts, integrating evidence across charts and time, and understanding how narrative organization and visual design communicate information. Yet existing benchmarks target either general videos or static charts, and data video understanding has not been systematically evaluated. We present DataVista, the first benchmark for data video understanding, containing 961 real-world data videos and 6,775 evaluation questions organized under a three-level progressive capability framework (data perception, temporal reasoning, narrative understanding) with 10 fine-grained question types across five topic domains. Systematic evaluation of 19 mainstream MLLMs shows that the best-performing model, Gemini-3.1-Pro, achieves 70.0% overall accuracy, still far below human expert performance, with models performing worst on Causal Reasoning and Narrative Structure. Increasing frame counts and adding subtitles mainly benefit data perception and temporal reasoning, with limited gains in narrative understanding. Further analysis of model responses identifies typical failure modes in chart reading, evidence judgment, and instruction understanding. The benchmark is available at https://github.com/HKUSTDial/DataVista.

Comments46 pages, 22 figures, 14 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑