Metric-Bench:探索室内场景中VLM的上下文空间度量推理
Metric-Bench: Exploring In-context Spatial Metric Reasoning in VLMs for Indoor Scenes
- State Key Lab of CAD & CG, Zhejiang University(浙江大学CAD&CG国家重点实验室)
- Zhejiang University of Technology(浙江工业大学)
- Hangzhou Dianzi University(杭州电子科技大学)
- Xiaomi Corporation(小米公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对VLM空间度量推理受限于像素级监督的问题,提出Metric-Bench基准和MetricReasoner微调方法,利用参考物体隐式学习2D-3D映射,显著提升度量理解与下游具身性能,且不损害通用能力。
AI中文摘要:
度量推理是视觉语言模型(VLMs)的一项关键且具有挑战性的任务,在具身智能任务(如机器人操作和自主导航)中发挥着核心作用。然而,当前的空间推理仍受限于僵化的像素级监督;这种局部优化往往损害通用多模态智能,导致性能下降或广泛推理能力的灾难性遗忘。为解决这些局限,我们引入了Metric-Bench,一个专注于利用上下文信息引导度量空间推理的基准。通过融入图像内具有已知物理尺寸的参考物体,Metric-Bench引导模型在无需相机内参的情况下隐式学习2D到3D的映射。我们进一步提出了MetricReasoner,一种针对参考接地度量推理的任务自适应强化微调方案,采用结构化提示和可验证的数值奖励。在Metric-Bench上的大量实验表明,我们的方法显著增强了空间度量理解,超越现有甚至更大的专有模型达43.1%,同时在下游具身任务上相较于空间专用模型在RoboSpatial总体准确率上提升30.4%,在ERQA上提升9.3%,并在通用基准上持续带来增益(V$\star$Bench上提升15.9%,BLINK上提升88.9%),表明所提出的自适应方法不一定损害通用VLM能力。
英文摘要:
Metric reasoning is a critical and challenging task for Vision Language Models (VLMs), playing a pivotal role in embodied AI tasks such as robotic manipulation and autonomous navigation. However, current spatial reasoning remains bottlenecked by rigid pixel-level supervision; such localized optimization often compromises general multimodal intelligence, triggering performance degradation or catastrophic forgetting of broad reasoning capabilities. To address these limitations, we introduce Metric-Bench, a focused benchmark designed to guide metric-spatial reasoning using contextual information. By incorporating in-image reference objects with known physical dimensions, Metric-Bench guides models to implicitly learn the 2D-to-3D mapping without camera intrinsics. We further present MetricReasoner, a task-adapted reinforcement fine-tuning recipe for reference-grounded metric reasoning, using structured prompts and verifiable numerical rewards. Extensive experiments on Metric-Bench demonstrate that our approach significantly enhances spatial metric understanding, outperforming existing and even larger proprietary models by 43.1\%, while improving downstream embodied performance over a spatial-specialized counterpart by 30.4\% on RoboSpatial overall accuracy and 9.3\% on ERQA, and additionally delivering consistent gains on general benchmarks (15.9\% on V$\star$Bench, 88.9\% on BLINK), indicating that the proposed adaptation does not necessarily compromise general VLM capabilities.