arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DINOcular:自监督视觉空间表征学习

DINOcular: Self-Supervised Visuospatial Representations

Farkhat Almukhamedov, Sami Azirar, Hermann Blum

arXiv 2608.27226首次发表:更新:

发表机构

University of Bonn(波恩大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DINOcular是一种从RGB-D观测学习联合视觉空间表征的自监督框架,通过块间与块内融合集成几何先验,在3D几何基准上优于同规模方法,且在RGB-D语义分割任务中保持竞争力。

AI 中文摘要

我们提出了一种从RGB-D观测中学习联合视觉空间表征的自监督框架。当前视觉基础模型几乎仅在RGB图像上训练,而许多具身系统可获取显式深度感知信息,提供单目输入无法恢复的几何信息。我们的方法通过块间与块内融合,将深度衍生的几何先验与视觉骨干网络集成,使模型能高效编码外观与空间结构。所得表征在保留语义迁移能力的同时,展现出3D感知方面的显著提升:在多个3D几何基准上,其性能优于同等规模的现有方法;在标准RGB-D语义分割任务测试中,仍保持竞争力。

英文摘要

We introduce a self-supervised framework for learning joint visuospatial representations from RGB-D observations. While modern vision foundation models are trained almost exclusively on RGB images, many embodied systems have access to explicit depth sensing, which provides geometric information that monocular inputs cannot recover. Our method integrates depth-derived geometric priors with a visual backbone through inter-patch and intra-patch fusion, enabling the model to encode both appearance and spatial structure efficiently. The resulting representation shows promising improvements on 3D awareness while preserving semantic transfer: it outperforms prior methods of comparable scale on multiple 3D geometry benchmarks, and remains competitive when probed for standard RGB-D semantic segmentation tasks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑