arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.15176cs.AIcs.CLcs.HC

用于科学可视化素养的多模态大语言模型基准测试

Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy

Patrick Phuoc Do, Chau M. Ta, Chaoli Wang

首次发表
浏览论文内容

中文总结 AI 辅助

研究对六个多模态大语言模型进行科学可视化素养基准测试,涵盖多种技术和任务类型。通过封闭世界协议评估闭源和开源模型,与人类参与者数据对比。发现模型表现不均,Gemini最强,开源模型低于人类基线,明确SciVis素养对评估多模态AI系统的必要性。

中文摘要 AI 辅助

多模态大语言模型(MLLMs)越来越多地用于解释可视化,但目前的评估主要以图表为中心,对科学可视化(SciVis)理解的证据有限。我们在科学可视化素养评估测试中对六个MLLMs进行基准测试,该测试是一项标准化的SciVis素养评估,包括基于18个科学可视化和插图的49个项目,涵盖8种技术和11种任务类型。我们在封闭世界协议下评估了三个闭源模型和三个开源模型,并使用485名人类参与者的数据比较了它们的性能。结果表明,当前的MLLMs没有表现出统一的SciVis素养。Gemini是总体上最强的模型,在评估子集中超过了人类平均水平,而开源模型仍低于人类基线。不同技术和任务的性能差异很大:模型在科学插图、搜索和空间理解方面表现最佳,但在基于纹理和基于集成的可视化以及定量估计方面存在困难。错误分析揭示了在细粒度定量估计、流向解释和基础编码解释方面反复出现的失败。这些发现将SciVis素养定位为评估多模态人工智能系统的必要基准维度。我们的代码和模型输出可在这个https URL上公开获取。

英文摘要

Multimodal large language models (MLLMs) are increasingly used to interpret visualizations, yet current evaluations remain largely chart-centric and provide limited evidence of understanding of scientific visualization (SciVis). We benchmark six MLLMs on the scientific visualization literacy assessment test, a standardized SciVis literacy assessment comprising 49 items based on 18 scientific visualizations and illustrations, spanning 8 techniques and 11 task types. We evaluate three closed-source and three open-source models under a closed-world protocol and compare their performance using data from 485 human participants. Results show that current MLLMs do not exhibit uniform SciVis literacy. Gemini is the strongest model overall, exceeding the human mean across the evaluated subsets, whereas the open-source models remain below the human baseline. Performance is highly uneven across techniques and tasks: models perform best on scientific illustration, search, and spatial understanding, but struggle on texture-based and integration-based visualizations and on quantitative estimation. Error analysis reveals recurring failures in fine-grained quantitative estimation, flow-direction interpretation, and grounded encoding interpretation. These findings position SciVis literacy as a necessary benchmark dimension for evaluating multimodal AI systems. Our code and model outputs are publicly available at https://github.com/patdmp/mllm-scivis-lit-benchmark.

发表机构

  • University of Notre Dame(圣母大学)

机构由 AI 辅助整理,请以论文原文为准。

↑