arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.25021cs.AIcs.HCcs.MAcs.SE

图表支持还是模型提供?审视多模态大语言模型生成的可访问可视化声明

Chart-Supported or Model-Supplied? Examining MLLM-Generated Claims for Accessible Visualization

Ishrat Jahan Eliza, Md Dilshadur Rahman

首次发表
浏览论文内容

中文总结 AI 辅助

研究多模态大语言模型生成可视化声明的证据基础,通过对多种来源、模型及输入条件的探索性研究,分析模型标签与数值一致性,发现可访问图表上下文有作用,添加图像及无上下文提示效果不佳,推动可区分证据支持声明与模型解释的描述系统发展。

中文摘要 AI 辅助

多模态大语言模型(MLLMs)能将可视化模式与外部因果及领域知识相联系,但其解释的证据基础常不明确。我们对来自四个来源的102个可视化、三个MLLMs以及四种输入条件进行了探索性研究,这些条件在图像访问、特定来源的可访问图表上下文和无上下文框架方面有所不同。在1224个描述中,我们分析了模型归因的直接、派生和推测性标签,并对数值一致性进行了自动审核。可访问图表上下文使Gemini和GPT倾向于直接声明,并提高了某些模型的数值一致性。将图像添加到完整上下文中并未产生一致的数值优势,无上下文提示也未可靠地增加谨慎语言。提示定义的现实世界意义部分主要仍是推测性的。这些结果推动了可访问描述系统的发展,该系统能区分由提供的证据支持的声明和模型提供的解释。

英文摘要

Multimodal large language models (MLLMs) can connect visualization patterns to external causes, consequences, and domain knowledge, but the evidential basis of these interpretations is often unclear. We present an exploratory study of 102 visualizations from four sources, three MLLMs, and four input conditions that vary access to the image, accessible chart context (non-image artifacts such as data tables, captions, alt text, and screen-reader structures), and withheld-context framing. Across 1,224 descriptions, we analyze model-attributed DIRECT, DERIVED, and SPECULATIVE labels and conduct an automated audit of numeric agreement. Accessible chart context shifted Gemini and GPT toward DIRECT claims and improved numeric agreement for some models. Adding the image to the full context did not yield a consistent numeric benefit, and the withheld-context prompt did not reliably increase cautious language. The prompt-defined Real-World Significance section remained predominantly SPECULATIVE. These results motivate accessible description systems that distinguish claims supported by supplied evidence from model-supplied interpretation.

补充信息

↑