arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25841cs.CVcs.MM

Metric-Bench:探索室内场景中VLM的上下文空间度量推理

Metric-Bench: Exploring In-context Spatial Metric Reasoning in VLMs for Indoor Scenes

  • State Key Lab of CAD & CG, Zhejiang University(浙江大学CAD&CG国家重点实验室)
  • Zhejiang University of Technology(浙江工业大学)
  • Hangzhou Dianzi University(杭州电子科技大学)
  • Xiaomi Corporation(小米公司)

机构由 AI 辅助整理,请以论文原文为准。

Yuling Xi, Haokai Zhang, Muzhi Zhu, Hao Zhong, Zongze Du, Hengyu Zhao, Chenchen Jing, Yufei Yin, Bin Qin, Yongjie Yang, Zhenbo Luo, Hao Chen, Chunhua Shen

AI总结:

针对VLM空间度量推理受限于像素级监督的问题,提出Metric-Bench基准和MetricReasoner微调方法,利用参考物体隐式学习2D-3D映射,显著提升度量理解与下游具身性能,且不损害通用能力。

AI中文摘要:

度量推理是视觉语言模型(VLMs)的一项关键且具有挑战性的任务,在具身智能任务(如机器人操作和自主导航)中发挥着核心作用。然而,当前的空间推理仍受限于僵化的像素级监督;这种局部优化往往损害通用多模态智能,导致性能下降或广泛推理能力的灾难性遗忘。为解决这些局限,我们引入了Metric-Bench,一个专注于利用上下文信息引导度量空间推理的基准。通过融入图像内具有已知物理尺寸的参考物体,Metric-Bench引导模型在无需相机内参的情况下隐式学习2D到3D的映射。我们进一步提出了MetricReasoner,一种针对参考接地度量推理的任务自适应强化微调方案,采用结构化提示和可验证的数值奖励。在Metric-Bench上的大量实验表明,我们的方法显著增强了空间度量理解,超越现有甚至更大的专有模型达43.1%,同时在下游具身任务上相较于空间专用模型在RoboSpatial总体准确率上提升30.4%,在ERQA上提升9.3%,并在通用基准上持续带来增益(V$\star$Bench上提升15.9%,BLINK上提升88.9%),表明所提出的自适应方法不一定损害通用VLM能力。

英文摘要:

Metric reasoning is a critical and challenging task for Vision Language Models (VLMs), playing a pivotal role in embodied AI tasks such as robotic manipulation and autonomous navigation. However, current spatial reasoning remains bottlenecked by rigid pixel-level supervision; such localized optimization often compromises general multimodal intelligence, triggering performance degradation or catastrophic forgetting of broad reasoning capabilities. To address these limitations, we introduce Metric-Bench, a focused benchmark designed to guide metric-spatial reasoning using contextual information. By incorporating in-image reference objects with known physical dimensions, Metric-Bench guides models to implicitly learn the 2D-to-3D mapping without camera intrinsics. We further present MetricReasoner, a task-adapted reinforcement fine-tuning recipe for reference-grounded metric reasoning, using structured prompts and verifiable numerical rewards. Extensive experiments on Metric-Bench demonstrate that our approach significantly enhances spatial metric understanding, outperforming existing and even larger proprietary models by 43.1\%, while improving downstream embodied performance over a spatial-specialized counterpart by 30.4\% on RoboSpatial overall accuracy and 9.3\% on ERQA, and additionally delivering consistent gains on general benchmarks (15.9\% on V$\star$Bench, 88.9\% on BLINK), indicating that the proposed adaptation does not necessarily compromise general VLM capabilities.

补充信息

↑