发表机构
The University of Hong Kong(香港大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
LongChart VQA是面向具备复杂多图表推理能力的MLLMs的综合基准,该基准含平均6.5张图像与31.2个问题,评估10个SOTA MLLMs发现其准确率随计算复杂性提升下降,为多图表推理研究指明方向。
AI 中文摘要
多模态大语言模型(MLLMs)正凭借扩展的上下文窗口和更强的推理能力快速发展,能够实现多图表理解与多步骤推理。随着MLLMs被应用于复杂智能体任务,这些能力愈发重要。然而,现有基准大多侧重单图表感知,简单的图表间关联不足以评估上述能力。为捕捉多图表复杂性,同时确保一致性与有效性,我们设计了由潜在图支撑的合成流程。基于该流程,我们推出了LongChart基准,其视觉问答(VQA)集平均包含6.5张图像与31.2个问题。我们评估了10个最先进的MLLMs,考察了影响性能的三个因素:推理模式、辅助工具以及对图像扰动的鲁棒性。结果显示,随着计算复杂性提升,MLLMs的准确率下降且差异显著,这为多图表推理的未来研究指明了方向。
英文摘要
Multimodal large language models (MLLMs) are rapidly evolving with expanded context windows and stronger reasoning capabilities, enabling multi-chart understanding and multi-step inference. These abilities are increasingly important as MLLMs are adopted in complex agentic tasks. However, existing benchmarks largely emphasize single-chart perception, while simple chart-to-chart connections are insufficient to evaluate these capabilities. To capture multi-chart complexity while ensuring consistency and validity, we design a synthesis pipeline supported by latent graphs. Building on this pipeline, we introduce LongChart, a benchmark whose VQA sets contain an average of 6.5 images and 31.2 questions. We evaluate 10 state-of-the-art MLLMs and examine three factors that influence performance: reasoning patterns, auxiliary tools, and robustness to image perturbations. Our results show that MLLM accuracy decreases and varies substantially as computational complexity increases, highlighting directions for future research in multi-chart reasoning.