arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PointVGGT:基于视觉几何基础先验的零样本多视图RGB-D点云配准

PointVGGT: Zero-Shot Multiview RGB-D Point Cloud Registration with Visual Geometry Foundation Priors

Haobo Jiang, Liang Yu, Jianmin Zheng

arXiv 2610.11612首次发表:更新:

发表机构

Nanyang Technological University; Alibaba Group(南洋理工大学; 阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出PointVGGT框架,采用“先基础后优化”范式,利用视觉几何基础模型VGGT实现零样本多视图RGB-D点云配准,在多类数据集上验证了其高精度与高效性。

AI 中文摘要

本文针对多视图RGB-D点云配准问题,旨在为无序RGB-D扫描数据估计全局刚性位姿,并将其对齐到度量一致的坐标系中。传统的“先配对后全局”范式存在配对配准局部优化、严重误差传播及计算负担高的问题。现有方法通常将RGB数据仅视为辅助匹配线索,忽略跨图像序列编码的整体几何先验(如相机位姿和3D模型)。本文提出PointVGGT,这是一个基于新型“先基础后优化”范式的零样本框架,系统利用视觉几何基础模型(如VGGT)作为计算骨干,实现鲁棒、无需训练的多视图RGB-D配准。在基础阶段,我们通过将基础模型的尺度模糊位姿预测与度量深度观测对齐,直接恢复度量一致的全局位姿(无需任何配对估计)。在优化阶段,我们引入高效的体素化空间哈希机制,利用基础模型生成的全局一致3D重建作为共享空间锚点,实现近线性时间的密集多视图对应。在此基础上,使用共轭梯度求解器执行基于IRLS的鲁棒仅运动束调整,以联合最小化多视图位姿优化的对应残差和重投影残差。在室内、以对象为中心、室外数据集上的大量实验验证了所提方法出色的零样本配准精度和计算效率。

英文摘要

This paper addresses multiview RGB-D point cloud registration, aiming to estimate global rigid poses for unordered RGB-D scans and align them in a metrically consistent coordinate frame. The conventional pairwise-then-global paradigm suffers from locally optimized pairwise registration, severe error propagation and high computational burden. In particular, existing methods typically treat RGB data as a mere auxiliary matching cue and overlook the holistic geometric priors (e.g., camera poses and 3D models) encoded across image sequences. This paper introduces PointVGGT, a zero-shot framework built upon a novel \emph{foundation-then-refinement} paradigm that systematically leverages visual geometry foundation models (e.g., VGGT) as the computational backbone for robust, training-free multiview RGB-D registration. In the foundation stage, we directly recover metrically consistent global poses (without any pairwise estimation) by grounding the scale-ambiguous pose predictions of the foundation model against metric depth observations. In the refinement stage, we introduce an efficient voxelized spatial hashing mechanism that exploits the globally coherent 3D reconstruction (induced by the foundation model) as a shared spatial anchor, enabling dense multiview correspondences in near-linear time. On top of this, an IRLS-based robust motion-only bundle adjustment is performed using a conjugate gradient solver to jointly minimize the correspondence and reprojection residuals for multiview pose refinement. Extensive experiments on indoor/object-centric/outdoor datasets verify the outstanding zero-shot registration accuracy and computational efficiency of our proposed method.

Comments19 Pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑