arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VGGT 对重叠的理解:探索共可见性的几何基础模型

What VGGT Knows About Overlap: Probing Geometric Foundation Models for Co-Visibility

Filippo Ziliotto, Luciano Serafini, Lamberto Ballan, Tommaso Campari

arXiv 2607.09503首次发表:更新:

发表机构

University of Padova; Fondazione Bruno Kessler (FBK)(帕多瓦大学; 布鲁诺·凯斯勒基金会)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究三维重建和机器人定位中共可见性问题,通过探索 VGGT 模型,发现其隐式编码共可见性及层间特性,引入 Co-VGGT 方法,该方法冻结 VGGT 并训练轻量级头,在基准测试中表现出色,提升了共可见性分类效果。

AI 中文摘要

三维重建和机器人定位中的一个基本挑战是共可见性,即在最小重叠场景中确定哪些图像对共享重叠可见表面。我们证明 VGGT 隐式地将共可见性编码为一种涌现行为,其内部表示呈现出类似于大语言模型的层次结构。我们识别出 L17 层为负锚点。在此基础上,我们引入 Co-VGGT,它冻结 VGGT 并仅训练一个轻量级的逐层专家混合头来仅从 RGB 分类共可见性。在 Co-VisiON 基准测试中,Co-VGGT 超过了人工标注基线,成对和多视图方面比先前工作有显著提升,成对预测校准良好,可直接用作下游 SfM 和 SLAM 管道中可见性图的边权重。

英文摘要

A fundamental challenge in 3D reconstruction and robotic localization is co-visibility: determining which image pairs share overlapping visible surfaces, particularly in scenarios with minimal overlap. We demonstrate that VGGT implicitly encodes co-visibility as an emergent behavior: without any supervision for this task, its internal representations exhibit a clear hierarchical structure mirroring that of large language models, i.e. early layers build a 3D-aware scene representation, while late layers act as dedicated co-visibility reasoners. In particular, we identify layer L17 as a negative anchor that consistently routes non-co-visible pairs for this backbone, regardless of the evaluation setting, providing task-grounded evidence of layer specialization in a geometry-grounded foundation model. Building on this, we introduce Co-VGGT, which freezes VGGT and trains only a lightweight layer-wise mixture-of-experts head (less than 7.5M parameters) to classify co-visibility from RGB alone, treating each layer as a specialized expert whose geometric abstraction is adaptively weighted per input pair. On the Co-VisiON benchmark, Co-VGGT surpasses the human annotation baseline and improves over prior work by more than 25% pairwise and 10% multiview. Pairwise predictions are well-calibrated (ECE=0.030), enabling direct use as edge weights in visibility graphs for downstream SfM and SLAM pipelines without post-hoc correction. Code and data are available.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑