GraphVerse:面向多模态大语言模型的综合性视觉图推理基准
GraphVerse: A Comprehensive Visual Graph Reasoning Benchmark for Multimodal Large Language Models
另 3 家 · 查看机构详情
- New York University(纽约大学)
- Sensetime Research(商汤科技研究院)
- Tsinghua University(清华大学)
- New York University Shanghai(上海纽约大学)
- University of Georgia(佐治亚大学)
- The Hong Kong Polytechnic University(香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对现有多模态大语言模型(MLLMs)评估缺乏结构化视觉推理测试的问题,提出GraphVerse基准,通过以图为中心的图像编辑(GIE)策略和VGR-Score指标,评估MLLMs在单/配对图像场景下的视觉图推理能力,揭示其相关局限并验证方法有效性。
中文摘要 AI 辅助
近期,多模态大语言模型(MLLMs)在各类视觉-语言任务中取得了显著进展,催生了对更具挑战性基准的迫切需求。然而现有评估仍难以深入洞察这些模型是否能真正对结构化视觉信息进行推理。视觉图推理(VGR)为应对这一挑战提供了极具吸引力的测试平台,要求模型整合感知、结构理解以及对基于图的视觉输入的多步推理。但现有VGR基准常将任务简化为视觉感知后接文本推理,仅限制在单图像场景下评估,依赖仅含答案的指标,且未能充分覆盖以图为中心的真实场景。为弥合这一差距,我们推出GraphVerse,这是一个统一的基准,在单图像和配对图像场景下共同评估MLLMs的感知、视觉推理和基于文本的图推理能力。其核心是一套以图为中心的图像编辑(GIE)策略,可修改图图像同时保留其语义,将其转化为视觉推理的主动测试。我们进一步提出VGR-Score,这是一种对过程敏感的指标,可评估超越最终答案准确率的推理质量。大量实验揭示了当前MLLMs在VGR中的若干关键局限,同时验证了GIE策略的有效性以及GraphVerse向更广泛多模态推理能力的可迁移性。代码可在该https链接获取。
英文摘要
Recent Multimodal Large Language Models (MLLMs) have achieved remarkable progress across diverse vision-language tasks, creating an urgent need for more challenging benchmarks. Yet existing evaluations still provide limited insight into whether these models can truly reason over structured visual information. Visual Graph Reasoning (VGR) offers a compelling testbed for this challenge, requiring models to integrate perception, structural understanding, and multi-step reasoning over graph-based visual inputs. However, prior VGR benchmarks often reduce the task to visual perception followed by text-based reasoning, restrict evaluation to single-image settings, rely on answer-only metrics, and underrepresent realistic graph-centric scenarios. To bridge the gap, we introduce GraphVerse, a unified benchmark that jointly evaluates perception, visual reasoning, and text-based graph reasoning in MLLMs under both single-image and paired-image settings. At its core is a suite of Graph-centric Image Editing (GIE) strategies that modify graph images while preserving their semantics, turning them into active tests of visual reasoning. We further propose VGR-Score, a process-sensitive metric that evaluates reasoning quality beyond final-answer accuracy. Extensive experiments reveal several key limitations of current MLLMs in VGR, while also validating the effectiveness of GIE strategies and the transferability of GraphVerse to broader multimodal reasoning capabilities. The code is available at https://github.com/sunyuanfu/GraphVerse.