发表机构
NVIDIA(英伟达)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现实杂乱环境中需操作遮挡物体的视觉问答问题,提出PROBE框架,含模拟器、基准测试集及微调方案,提升VLM智能体性能并验证模拟到现实的迁移。
AI 中文摘要
视觉语言模型(VLM)在静态场景的二维定位、空间推理以及基于智能体工具的规划方面表现出色。但试想向家用机器人提问“我的药物还在柜子里吗?”,答案可能被一排容器物理遮挡,需先移开容器才能获取。在现实杂乱环境中回答此类问题需要动态场景推理:必须操作干扰项以揭示被遮挡物体,且每次动作都会改变模型需推理的场景。我们将此场景形式化为基于操作的视觉问答(MG-VQA),并引入PROBE——一个用于在此类任务上对VLM智能体进行基准测试和微调的框架。我们首先开发PROBE-Sim,一个具有日常物体和配备抓取、推动工具的机械臂的高保真桌面模拟器。PROBE-Sim用于创建PROBE-Bench:一个包含6类问题的150项任务的评估套件,适用于杂乱桌面场景,其中VLM需感知、拿起或推动物体后再回答问题。我们观察到所有前沿VLM都存在一致趋势:基于智能体工具的方法在所有任务类型上比仅感知的基准方法表现更好(平均提升8.0%)。我们进一步设计PROBE-Agent,一种微调方案,通过混合数据配方从强大的教师基础模型中提炼成功轨迹到较小的开放权重模型,该配方鼓励高效操作的问答。经PROBE-Agent微调的模型比其现成智能体基准表现更好(平均提升11.5%),并展示了对未见过的物体和保留任务的正迁移。我们通过在现实桌面环境中部署PROBE-Agent微调策略验证了模拟到现实的迁移。
英文摘要
Vision-language models (VLMs) have shown promising spatial reasoning capabilities from static visual inputs, where the evidence needed to answer a question is available in the provided views. However, in cluttered environments, answer-relevant evidence may be occluded rather than absent: an object may lie beneath a pile, be covered by another object, or have identifying information facing away from the camera. Answering such questions requires physical interaction to reveal the hidden evidence. We formalize this setting as Manipulation-Grounded Visual Question Answering (MG-VQA), where an agent answers a question about an initially cluttered scene by using manipulation as an intermediate evidence-gathering operation. We introduce MG-VQA-Bench, comprising 600 human-verified questions across four spatial reasoning tasks, and evaluate it in a cluttered tabletop interactive environment. Eight VLMs are evaluated under three levels of environment access: Direct (image only), Perception (pointing, segmentation, and scene graphs), and Manipulation (perception, grasping, and pushing). Across all models, Direct and Perception achieve average success rates of 36.2% and 37.6%, respectively, only slightly above an image-blind chance baseline of 32.7%. In contrast, Manipulation raises average success to 56.8% (40.8-83.0% across models). Stronger tool-calling VLMs, GPT~6~Astra (83.0%) and Gemini~3.8~Flash (69.3%), search persistently and re-ground after unsuccessful interactions, while weaker models often answer without gathering sufficient evidence or stop prematurely after failed actions. Our results highlight the need for VLM agents that use manipulation for persistent, physically grounded evidence gathering and recovery. Code and benchmark is available here: https://github.com/vineet2104/MG-VQA
CommentsThe previous version of this paper (V1) was called: "Probe: Manipulation-Grounded Visual Question Answering with VLM Agents"