发表机构
Shanghai Artificial Intelligence Laboratory; Zhejiang University(上海人工智能实验室; 浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LVLMs文档理解中视觉推理基准测试不足的问题,提出OmniMapBench,含2096个问答对,设计了视觉依赖指数(VDI)量化基准属性,评估25个LVLMs发现性能差距大,推动了以视觉为中心的推理进展。
AI 中文摘要
近期语言与视觉基础模型(LVLMs)的进展需要针对复杂的、基于视觉的推理建立强大的基准测试。许多文档理解基准测试存在关键局限:视觉内容常可简化为文本,无需真正的视觉基础就能实现高性能。为解决此局限,引入OmniMapBench以促进地图文档的以视觉为中心的推理。该基准测试包含来自九个类别的1603个地图文档中的2096个手动注释的问答对,旨在探究从感知到多步视觉推理的技能层次。为量化基准测试属性,提出了一种简单而有效的基准测试级度量:视觉依赖指数(VDI),即图像被与问题无关的描述替换时的准确率下降。OmniMapBench的VDI高于现有基准测试,定量验证了其对不可简化的视觉推理的关注。对25个领先的LVLMs在OmniMapBench上进行了全面评估,观察到显著的性能差距,表现最佳的模型准确率仅为75.03%。这一结果凸显了OmniMapBench对当前LVLMs构成的挑战。这项工作旨在推动LVLMs文档理解中以视觉为中心的推理取得进展。数据集和代码可通过此https URL公开获取。
英文摘要
Recent advancements in LVLMs necessitate robust benchmarks for complex, visually grounded reasoning. A critical limitation is identified in many document understanding benchmarks: visual content is often reducible to text, enabling high performance without genuine visual grounding. To address this limitation, OmniMapBench is introduced to foster visual-centric reasoning for map documents. The benchmark comprises 2,096 manually annotated question-answer pairs across 1,603 map documents from nine categories. It is designed to probe a hierarchy of skills, ranging from perception to multi-step visual reasoning. To quantify benchmark properties, a simple yet effective benchmark-level metric is proposed: the Visual Dependency Index (VDI), defined as the accuracy drop when images are replaced with question-agnostic descriptions. OmniMapBench exhibits higher VDI than established benchmarks, which quantitatively validates its focus on irreducible visual reasoning. Comprehensive evaluations of 25 leading LVLMs are conducted on OmniMapBench. A significant performance gap is observed, with the top-performing model achieving only 75.03\% accuracy. This result underscores the challenges posed by OmniMapBench to current LVLMs. This work aims to catalyze progress in visual-centric reasoning for document understanding of LVLMs. The dataset and code are publicly available at https://github.com/SIGMME/OmniMapBench.