arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39227cs.CVcs.AI

通过自蒸馏涌现的多视图几何

Emergent Multi-View Geometry Through Self-Distillation

  • Chalmers University of Technology(查尔姆斯理工大学)
  • LIGM, Ecole des Ponts, Univ. Gustave Eiffel, CNRS(LIGM,巴黎高科路桥学校,古斯塔夫·埃菲尔大学,法国国家科学研究中心)
  • Linköping University(林雪平大学)
  • Univ. Bordeaux, CNRS, Bordeaux INP, IMS, UMR 5218(波尔多大学,法国国家科学研究中心,波尔多国立理工学院,IMS实验室,UMR 5218)

机构由 AI 辅助整理,请以论文原文为准。

David Nordström, Thibaut Loiseau, Vincent Lepetit, Michael Felsberg, Guillaume Bourmaud, Fredrik Kahl

AI总结:

提出自监督方法Poincar3,通过自蒸馏从多视图学习表征,无需RGB重建或3D监督,在对应、位姿和3D重建任务上超越现有方法,并更准确编码相机运动。

AI中文摘要:

一个多世纪前,亨利·庞加莱提出,静止的观察者无法获得空间概念。然而,大多数视觉表征学习方法处理的是单张图像,而那些利用多视图的方法则依赖RGB重建,将几何与外观纠缠在一起。我们提出Poincar3,一种自监督方法,通过自蒸馏而非RGB重建从多个视图学习表征。我们将掩码补丁和图像级蒸馏与一个观察额外视图的教师模型相结合,从而无需显式3D监督即可从头训练。Poincar3在对应估计、相机位姿估计和3D重建任务上优于之前的单视图和多视图自监督方法,如DINOv3、MuM和Muskie。使用轻量级庞加莱适配器,我们还发现所学特征比现有自监督表征更准确地编码了相机运动。

英文摘要:

Over a century ago, Henri Poincaré argued that a motionless observer cannot acquire the notion of space. Yet, most visual representation learning methods operate on individual images, while those that leverage multiple views rely on RGB reconstruction, entangling geometry with appearance. We propose Poincar3, a self-supervised method that learns representations from multiple views through self-distillation instead of RGB reconstruction. We combine masked patch and image-level distillation with a teacher that observes additional views, enabling training from scratch without explicit 3D supervision. Poincar3 outperforms both previous single and multi-view self-supervised approaches such as DINOv3, MuM, and Muskie on correspondence estimation, camera pose estimation, and 3D reconstruction. Using a lightweight Poincaré adapter, we also find that our learned features encode camera motion more accurately than existing self-supervised representations.

↑