VLMs 能描述,但不能测量:面向机器人操作的目标中心场景理解
VLMs Can Describe, But Not Measure: Object-Centric Scene Understanding for Robotic Manipulation
- University of Trento(特伦托大学)
- Mondragón University(蒙德拉贡大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本文提出一种VLM驱动的模块化感知框架,通过对象级分割、语义标注与深度结合,在151个桌面场景上显著提升定位与深度估计,并集成于机器人任务规划。
中文摘要 AI 辅助
在未见过的环境中进行机器人操作,既需要语义理解,也需要可靠的度量信息。虽然视觉-语言模型(VLMs)提供了强大的语义能力,但其几何估计仍然不够可靠。在本文中,我们提出了一种基于VLM驱动的模块化感知框架,利用现成方法进行场景理解。从单张RGB-D观测出发,场景被分割为对象级区域,由VLM进行标注,并与深度信息相结合,构建与任务无关的目标中心表示。在151个桌面场景上的实验表明,所提出的分解方法在保持强语义性能的同时,相较于直接VLM推理,显著提升了定位和深度估计的准确性。所得到的表示还集成到任务规划框架中,用于机器人执行。
英文摘要
Robotic operation in previously unseen environments requires both semantic understanding and reliable metric information. While vision--language models (VLMs) provide strong semantic capabilities, their geometric estimates remain less reliable. In this paper, we propose a VLM-driven, modular perception framework for scene understanding using off-the-shelf approaches. Starting from a single RGB-D observation, the scene is segmented into object-level regions, annotated by a VLM, and grounded with depth information to construct a task-independent object-centric representation. Experiments on 151 tabletop scenes show that the proposed decomposition preserves strong semantic performance while substantially improving localization and depth estimation over direct VLM inference. The resulting representation is also integrated with a task-planning framework for robotic execution.