arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38285cs.CVcs.AI

GaugeVLM:通过测量几何干预构建空间监督结构

GaugeVLM: Structuring Spatial Supervision with Measured Geometric Interventions

发表机构中国科学院自动化研究所模式识别国家重点实验室与多模态人工智能系统实验室 · 中国科学院大学人工智能学院 · 中国科学院大学前沿交叉科学学院
另 2 家 · 查看机构详情
  • NLPR & MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所模式识别国家重点实验室与多模态人工智能系统实验室)
  • School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
  • School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(中国科学院大学前沿交叉科学学院)
  • National University of Singapore(新加坡国立大学)
  • The Chinese University of Hong Kong(香港中文大学)

机构由 AI 辅助整理,请以论文原文为准。

Hongbo Wang, Zihan Lin, Wenkui Yang, Shiran Ge, Yuang Ai, Jie Cao, Huaibo Huang, Ran He

首次发表
浏览论文内容

中文总结 AI 辅助

GaugeVLM通过3D场景中的受控干预和GaugeDPO目标,将空间关系误差转化为偏好信号,显著提升VLM空间推理能力并泛化至自动驾驶等领域。

中文摘要 AI 辅助

视觉语言模型(VLM)在同一空间关系的不同视角下可能自相矛盾,并且当该关系发生变化时无法做出响应。解决这些失败需要能够捕捉观测之间误差幅度和几何依赖性的监督,而这两者在基于单个答案或序数偏好的训练中仍然是隐性的。因此,我们引入了GaugeVLM,它通过在显式3D场景中进行受控的对象和相机干预来使这种结构显式化,从而产生具有空间关系之间测量差异和跨视角共享真相的关联观测。为了将这种结构转化为学习信号,其核心目标GaugeDPO将测量误差转换为偏好边际,直接监督跨视角的规范排序,并将干预引起的答案几率对比与具有视角特定尺度的测量关系变化联系起来。我们的分析界定了规范预测误差,并确立了跨视角和干预约束可以同时满足。实验上,GaugeVLM在三个VLM骨干网络上相对于监督微调提高了所有10个已建立的空间指标,其中主要的7B模型在MSMU距离和QSpatial+上分别获得了15.0和18.9个百分点的提升。这些提升还扩展到自动驾驶和具身推理,展示了跨领域的稳健泛化能力。

英文摘要

Vision-language models (VLMs) can contradict themselves across views of the same spatial relation and fail to respond when that relation changes. Addressing these failures requires supervision that captures error magnitude and geometric dependencies across observations, both of which remain implicit in training on individual answers or ordinal preferences. Therefore, we introduce GaugeVLM, which makes this structure explicit through controlled object and camera interventions in explicit 3D scenes, producing linked observations with measured differences between spatial relations and shared truths across views. To translate this structure into learning signals, its core objective, GaugeDPO, converts measured errors into preference margins, directly supervises correct canonical rankings across views, and links intervention-induced answer-odds contrasts to measured relation changes with view-specific scales. Our analysis bounds canonical prediction error and establishes that the cross-view and intervention constraints can be jointly satisfied. Empirically, GaugeVLM improves all 10 established spatial metrics over supervised fine-tuning across three VLM backbones, with the main 7B model gaining 15.0 and 18.9 percentage points on MSMU distance and QSpatial+, respectively. These gains also extend to autonomous driving and embodied reasoning, demonstrating the robust generalization across domains.

↑