arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉-语言模型中空间推理的几何编码

Geometric Encoding for Spatial Reasoning in Vision-Language Models

Antonio Jun, Haoshui Yu, Zhengyi Lu, Huirong Fu, Yao Qiang

arXiv 2609.34148首次发表:更新:

发表机构

Hunter College; New York University; Oakland University(亨特学院; 纽约大学; 奥克兰大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出几何代码流水线,从视频计算显式空间结构并注入视觉-语言模型提示,无需训练,在VSI-Bench上平均准确率提升4.1点,绝对距离任务提升24.1点。

AI 中文摘要

视觉-语言模型(VLMs)在识别视频中出现的内容方面远比在推理其空间和时间属性(如度量距离、物体尺寸以及跨帧一致的物体身份)方面可靠得多。我们提出了几何代码(Geometric Code),这是一种从感知到几何的流水线,它从视频中计算显式的空间结构,并将其作为上下文提供给VLMs以增强推理能力。感知层分割并分类物体,并从单目RGB视频中恢复深度、相机姿态和内部参数。随后,一个确定性的几何引擎将这些输出进行反投影、合并和清理,形成空间代码,包括每个物体的位置、尺寸、数量、物体间距离、出现顺序和房间几何形状。该代码被序列化到VLMs的提示中,要么与视频一起提供,要么完全替代视频。具体而言,我们的方法中没有任何组件经过训练或微调。在VSI-Bench上,将2B和4B开源模型与空间代码增强后,平均准确率比仅使用帧的基线提高了+4.1个百分点,其中在数值估计任务(如绝对距离)上提升最大,达到+24.1个百分点。结果表明,通过语言通道传递的显式计算几何,能够恢复小型VLMs无法仅从像素中提取的空间能力。

英文摘要

Vision-Language Models (VLMs) are far more reliable at recognizing what appears in a video than at reasoning about its spatial and temporal properties, such as metric distances, object dimensions, and consistent object identities across frames. We present Geometric Code, a perception-to-geometry pipeline that computes explicit spatial structure from video and supplies it to VLMs as context to augment reasoning. A perception layer segments and classifies objects and recovers depth, camera pose, and intrinsics from monocular RGB video. A deterministic geometric engine then back-projects, merges, and cleans these outputs into a spatial code, including per-object positions, dimensions, counts, inter-object distances, appearance order, and room geometry. The code is serialized into VLMs' prompts, either alongside the video or replacing it entirely. Specifically, there is no component trained or fine-tuned in our approach. On VSI-Bench, augmenting 2B and 4B open models with the spatial code improves average accuracy by +4.1 points over the frames-only baseline, with the largest gains on numeric estimation tasks such as absolute distance (+24.1 points). The results suggest that explicitly computed geometry, delivered through the language channel, recovers spatial competence that small VLMs cannot extract from pixels alone.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑