arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

重新审视跨视图补全:基于重构误差对比的自监督预训练

Revisiting Cross-View Completion: Self-Supervised Pre-Training via Reconstruction Error Comparison

Thibaut Loiseau, Guillaume Bourmaud, Vincent Lepetit

arXiv 2609.01530首次发表:更新:

发表机构

LIGM, Ecole Nationale des Ponts et Chaussées, IP Paris, Univ Gustave Eiffel, CNRS; Univ. Bordeaux, CNRS, Bordeaux INP, IMS(LIGM实验室、巴黎国立路桥学校、巴黎IP大学、古斯塔夫·埃菲尔大学、法国国家科学研究中心; 波尔多大学、法国国家科学研究中心、波尔多国立综合理工学院、IMS研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出Gekko模型,通过重构误差对比实现跨视图补全的自监督预训练,无需3D标注即可提升多3D视觉任务性能,且可直接从原始视频训练,效果优于现有方法。

AI 中文摘要

跨视图补全的自监督预训练可从图像对的共可见区域学习3D视觉的强特征。然而,参考视图几乎无法为非共可见块的重构提供信息,在这些区域隐式产生单目训练信号。我们提出Gekko模型,将这一限制转化为有用信号:跨视图重构误差相对于掩码自编码器误差的相对提升可作为共可见性的自监督代理,大幅提升对应共可见区域,几乎无提升对应非共可见区域。Gekko是从头训练的网络,可联合执行跨视图补全、掩码自编码及该相对提升的逐像素预测,无需任何真实3D标注即可为所有掩码区域提供额外双目信号。在相同架构与训练数据下,Gekko在零样本对应估计、相对位姿估计、点云图回归任务中均持续优于CroCo;在最严格的相对位姿阈值下精度最高提升6倍,ETH3D数据集上的端点误差下降22%。其学习的额外通道本身是未见场景的强共可见性检测器,冻结后的Gekko特征优于同等或更大规模的已发布跨视图骨干网络。它还可通过简单的基于步长的课程直接从原始视频训练,消除了先前方法所需的繁琐3D预处理,同时与经精心整理数据训练的模型表现相当。代码与预训练模型已公开可用。

英文摘要

Self-supervised pre-training via cross-view completion learns strong features for 3D vision from co-visible regions of image pairs. However, the reference view provides little information for reconstructing non-co-visible patches, implicitly yielding a monocular training signal in these regions. We introduce Gekko, which turns this limitation into a useful signal. The relative improvement of the cross-view reconstruction error over a masked-autoencoder error is a self-supervised proxy for co-visibility: large improvements indicate co-visible regions, negligible ones non-co-visible areas. Gekko is a network, trained from scratch, that jointly performs cross-view completion, masked autoencoding, and per-pixel prediction of this relative improvement, providing an additional binocular signal for all masked regions without any ground-truth 3D annotation. Under identical architectures and training data, Gekko consistently outperforms CroCo on zero-shot correspondence estimation, relative pose estimation, and pointmap regression, with up to 6 times higher accuracy at the strictest relative-pose threshold and a 22% drop in end-point error on ETH3D. The extra channel it learns is itself a strong co-visibility detector on unseen scenes, and Gekko's frozen features outperform released cross-view backbones of comparable or larger size. It can also be trained directly from raw videos with a simple stride-based curriculum, removing the cumbersome 3D preprocessing prior methods require while matching models trained on curated data. Code and pre-trained models are publicly available.

CommentsProject page: https://thibautloiseau.github.io/projects/gekko

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑