arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SeeSE3:视觉特征中3D空间的出现

SeeSE3: Emergence of 3D Space in Vision Features

Caroline Chen, Sayna Ebrahimi, Fedor Kitashov, Ming-Hsuan Yang, Leonidas Guibas, Viorica Pătrăucean, Maks Ovsjanikov

arXiv 2607.14228首次发表:更新:

发表机构

Google DeepMind(谷歌DeepMind)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究视觉基础模型构建的表征与3D欧几里得空间内在属性的关系,提出从拓扑和几何角度评估的探测器,发现自监督视觉模型有相关潜在子空间,并基于此提出“潜在空间导航”技术用于视觉里程计和定位。

AI 中文摘要

在本文中,我们探讨视觉基础模型构建的表征是否反映3D欧几里得空间的内在属性。与以往通过回归深度或法线等以图像为中心的量来探究视觉特征的3D感知的工作不同,我们研究视觉特征空间结构与欧几里得变换群$SE(3)$之间的关系。我们提出了一组从拓扑和几何角度评估这种关系的探测器:一个衡量特征邻域与空间拓扑之间对齐的相互邻域度量,以及一个用于测试静态场景中潜在位移对相机运动几何的线性可达性的庞加莱适配器。我们表明,原则上未经过直接3D监督或主动代理训练的自监督视觉模型,在正确探测时,拥有与三维欧几里得空间显著高度相关的潜在子空间。基于这一见解,我们提出了一类新的“潜在空间导航”技术,该技术可在潜在空间中纯执行视觉里程计和定位,无需显式3D重建。

英文摘要

In this paper, we ask whether vision foundation models construct representations that reflect the intrinsic properties of 3D Euclidean space. Unlike previous works that probe 3D awareness of vision features by regressing image-centric quantities such as depth or normals, we investigate the relation between the structure of the space of visual features and the group of Euclidean transformations $SE(3)$. We propose a set of probes to evaluate this relation from both topological and geometric perspectives: a mutual neighborhood metric that measures the alignment between feature neighborhoods and spatial topology, and a Poincaré Adapter to test the linear accessibility of the geometry of camera motion from latent displacements in static scenes. We show that self-supervised vision models, which, in principle, have not been trained with direct 3D supervision or active agency, possess latent subspaces that are remarkably strongly correlated with three-dimensional Euclidean space, when probed correctly. Building on this insight we propose a new class of "Latent-Space Navigation" techniques that perform visual odometry and localization purely in the latent space, bypassing the need for explicit 3D reconstruction.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑