arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LongChart VQA:面向具备复杂多图表推理能力的多模态大语言模型的综合基准

LongChart VQA: A Comprehensive Benchmark for MLLMs with Complex Multi-Chart Reasoning

Ziyan Xiao, Yinghao Zhu, Wenting Zhang, Heaju Kim, Lequan Yu

arXiv 2608.01328首次发表:更新:

发表机构

The University of Hong Kong(香港大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

LongChart VQA是面向具备复杂多图表推理能力的MLLMs的综合基准,该基准含平均6.5张图像与31.2个问题,评估10个SOTA MLLMs发现其准确率随计算复杂性提升下降,为多图表推理研究指明方向。

AI 中文摘要

多模态大语言模型(MLLMs)正凭借扩展的上下文窗口和更强的推理能力快速发展,能够实现多图表理解与多步骤推理。随着MLLMs被应用于复杂智能体任务,这些能力愈发重要。然而,现有基准大多侧重单图表感知,简单的图表间关联不足以评估上述能力。为捕捉多图表复杂性,同时确保一致性与有效性,我们设计了由潜在图支撑的合成流程。基于该流程,我们推出了LongChart基准,其视觉问答(VQA)集平均包含6.5张图像与31.2个问题。我们评估了10个最先进的MLLMs,考察了影响性能的三个因素:推理模式、辅助工具以及对图像扰动的鲁棒性。结果显示,随着计算复杂性提升,MLLMs的准确率下降且差异显著,这为多图表推理的未来研究指明了方向。

英文摘要

Multimodal large language models (MLLMs) are rapidly evolving with expanded context windows and stronger reasoning capabilities, enabling multi-chart understanding and multi-step inference. These abilities are increasingly important as MLLMs are adopted in complex agentic tasks. However, existing benchmarks largely emphasize single-chart perception, while simple chart-to-chart connections are insufficient to evaluate these capabilities. To capture multi-chart complexity while ensuring consistency and validity, we design a synthesis pipeline supported by latent graphs. Building on this pipeline, we introduce LongChart, a benchmark whose VQA sets contain an average of 6.5 images and 31.2 questions. We evaluate 10 state-of-the-art MLLMs and examine three factors that influence performance: reasoning patterns, auxiliary tools, and robustness to image perturbations. Our results show that MLLM accuracy decreases and varies substantially as computational complexity increases, highlighting directions for future research in multi-chart reasoning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑