发表机构
Google(谷歌)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对VLMs多视图集成能力评估不足问题,引入MultiView-Bench基准,发现模型在3D空间关系等方面的局限及偏差,提出ViewNavigator框架,能主动选择视角、融合证据,提升基础模型在该基准上的表现。
AI 中文摘要
近期视觉语言模型(VLMs)的基准测试大多评估单视图或有限视图感知,未测试将跨视角观察整合为连贯的、以世界为中心(非中心)3D心理模型的核心认知能力。我们引入了MultiView-Bench,这是一个专门设计用于评估整体3D场景理解的多视图集成的诊断基准。与现有专注于像素级映射或相机相对导航的数据集不同,MultiView-Bench要求模型将对象定位与瞬态视角解耦,并将其置于固定的全局坐标系中。这种能力是VLMs在部署到下游任务(如机械零件组装)之前的先决条件。我们对前沿VLMs的系统评估揭示了一致的失败模式:在单图像的2D平面关系上表现强劲,但在3D空间关系和跨视图聚合信息方面存在明显困难。我们还发现了VLMs中的偏差,如在非常规轴方向上的困难以及对对象颜色和纹理变化的敏感性。认识到这些局限性,我们提出了ViewNavigator,这是一个多智能体框架,它主动选择信息丰富的视角,感知并融合多视图证据,即使在严格的预算匹配比较下,也能在MultiView-Bench上改进各种基础模型(对于完整智能体提高3至5倍)。
英文摘要
Recent benchmarks for VLMs largely assess single- or limited-view perception, leaving untested the core cognitive ability to integrate observations across viewpoints into a coherent, world-centric (allocentric) 3D mental model. We introduce MultiView-Bench, a diagnostic benchmark expressly designed to evaluate multi-view integration for holistic 3D scene comprehension. Unlike existing datasets that focus on pixel-level mapping or camera-relative navigation, MultiView-Bench requires models to decouple object positioning from transient perspectives and ground them in a fixed global coordinate system. This capability serves as a prerequisite for VLMs before being deployed for downstream tasks such as mechanical part assembly. Our systematic evaluation of frontier VLMs reveals consistent failure modes: strong performance on 2D planar relations from a single image, but marked difficulty with 3D spatial relations and with aggregating information across views. We further identify biases in VLMs, such as struggles with unconventional axis directions and sensitivity to object colorways and texture variations. Acknowledging these limitations, we propose ViewNavigator, which uses active viewpoint selection and evidence fusion to improve four base models by 12.3--20.0 percentage points under a six-image cap matching the fixed-view baseline; budget-extended gains are model-dependent and reach 27 percentage points for GPT-5.