arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ChartDensity-Bench:在视觉密度下对多模态大语言模型进行数值数据重建的基准测试

ChartDensity-Bench: Benchmarking MLLMs for Numerical Data Reconstruction under Visual Density

Xinhe Wu, Yadong Jin

arXiv 2609.38781首次发表:更新:

发表机构

iSoftStone AI Research Lab (ISSAIR); iSoftStone(软通动力AI研究实验室; 软通动力)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出ChartDensity-Bench基准,通过控制图表密度评估多模态大语言模型从复合图表中重建数值数据的能力,发现密度增加导致重建性能普遍下降,且不同模型退化程度差异显著。

AI 中文摘要

多模态大语言模型(MLLMs)为从科学图表中恢复数值数据提供了一种有前景的方法,但它们在视觉密集图形中重建图表数据的能力仍知之甚少。现有的图表理解基准主要评估问答或图表级推理,对从科学图形中评估结构化数值重建的支持有限。我们引入了\textbf{ChartDensity-Bench},这是一个在受控视觉密度下评估MLLMs从复合图表图形中进行结构化数值数据重建的基准。该基准基于与源级真实数据配对的图表构建,系统地变化同时呈现的图表数量($k\in{1,3,6,9}$),从而实现对密度引起的性能下降的受控评估。我们进一步提出了一个多维评估框架,涵盖结构可靠性、重建完整性、可解析性和数值保真度。对五个近期MLLMs的实验表明,随着视觉密度的增加,数值重建通常会退化,而退化的程度在不同模型间差异显著。图表级配对比较进一步表明,当同一源图表嵌入更密集的视觉上下文时,其重建误差会更高。这些发现凸显了视觉密度是MLLM图表数据重建中一个重要的且此前未被充分探索的因素,并为评估模型在此设置下的鲁棒性提供了一个系统基准。

英文摘要

Multimodal large language models (MLLMs) offer a promising approach for recovering numerical data from scientific charts, but their ability to reconstruct chart data from visually dense figures remains poorly understood. Existing chart understanding benchmarks primarily evaluate question answering or chart-level reasoning and provide limited support for evaluating structured numerical reconstruction from scientific figures. We introduce \textbf{ChartDensity-Bench}, a benchmark for evaluating MLLMs on structured numerical data reconstruction from compound chart figures under controlled visual density. Built from charts paired with source-level ground-truth data, ChartDensity-Bench systematically varies the number of simultaneously presented charts ($k\in{1,3,6,9}$), enabling controlled evaluation of density-induced degradation. We further propose a multi-dimensional evaluation framework covering structural reliability, reconstruction completeness, parseability, and numerical fidelity. Experiments on five recent MLLMs show that numerical reconstruction generally degrades as visual density increases, while the magnitude of degradation varies substantially across models. Chart-level paired comparisons further show that the same source chart can incur higher reconstruction error when embedded in denser visual contexts. These findings highlight visual density as an important and previously underexplored factor in MLLM chart data reconstruction and provide a systematic benchmark for evaluating model robustness in this setting.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑