发表机构
Fudan University; The Hong Kong University of Science and Technology (Guangzhou)(复旦大学; 香港科技大学(广州))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究推出DVBench基准,评估MLLMs在结合动态图表与叙事的数据视频理解上的表现,发现开源模型性能不严格随参数规模扩展等现象,为MLLM发展提供参考。
AI 中文摘要
尽管多模态大语言模型(MLLMs)在图表理解和视频理解方面已取得显著进展,但当前的评估大多将这些能力孤立开来,导致在理解时间演化的结构化视觉信息方面存在关键缺口。为解决这一缺口,我们推出了DVBench,这是一个用于评估MLLMs在数据视频上表现的基准,数据视频是一种将动态图表与结构化叙事相结合的叙事媒介。我们将数据视频理解分解为五个维度。DVBench包含300个真实世界数据视频和1000个人工验证的问答对,这些问答对通过严格的半自动化流程筛选得到。对9个MLLMs的广泛评估显示,Gemini-3.1-Pro取得了最佳整体性能,而Kimi-k2.5是最强的开源模型。我们进一步发现两个值得注意的现象:开源模型的性能并不严格随参数规模扩展,且叙事能力并不能保证视觉能力。细粒度分析和消融研究进一步揭示了各维度的特定弱点,以及帧配置和字幕输入的影响,为未来MLLM的发展提供了参考。DVBench可通过此httpsURL公开获取。
英文摘要
While MLLMs have made significant strides in chart comprehension and video understanding, current evaluations largely isolate these capabilities, leaving a critical gap in understanding temporally evolving structured visual information. To address this gap, we introduce DVBench, a benchmark for evaluating MLLMs on data videos, a storytelling medium that integrates dynamic charts with structured narratives. We decompose data video understanding into five dimensions. DVBench comprises 300 real-world data videos and 1,000 human-verified QA pairs curated through a rigorous semi-automated pipeline. Extensive evaluations of nine MLLMs show that Gemini-3.1-Pro achieves the best overall performance, while Kimi-k2.5 is the strongest open-source model. We further identify two notable phenomena: open-source model performance does not scale strictly with parameter size, and narrative proficiency does not guarantee visual capability. Fine-grained analyses and ablation studies further reveal dimension-specific weaknesses and the effects of frame configurations and subtitle inputs, informing future MLLM development. DVBench is publicly available at https://bomiaowang.github.io/DVBench/.