发表机构
Chung-Ang University; Kyung Hee University(中央大学; 庆熙大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对3D视觉基础模型测试时几何一致性不足的问题,提出即插即用的Self-Geometry自适应流水线,结合多视图与对极一致性损失等,在6种模型和4个基准上实现位姿与几何估计的一致提升。
AI 中文摘要
近期视觉基础模型(VFMs)可通过单次前向传播预测深度、相机位姿和点云图,无需针对单场景优化,已具备较强泛化能力。然而,在VFMs预训练过程中未施加显式多视图几何一致性约束(如通过光束平差法实现),因此会出现不一致问题,且该约束的计算成本较高。现有工作在测试时强制执行从模型输出(如点云图、特征)衍生的隐式自一致性约束,但性能提升有限,尤其在预训练VFMs精度极低的场景中表现更差。与上述隐式信号不同,本文提出Self-Geometry,这是一种即插即用的测试时自适应流水线,直接以2D像素对应关系为伪真值施加显式多视图几何约束。Self-Geometry包含三部分:几何解耦优化,其结合多视图一致性与对极一致性损失,并通过梯度解耦防止梯度冲突;帧角度邻域,一种基于SO(3)测地距离的视图采样器,用于适度施加上述约束;轻量级TTA,通过LoRA对VFMs进行自适应调整。本文方法在6种VFMs(VGGT、π³、DA3-Giant/Large/Base/Small)和4个基准数据集(7Scenes、ETH3D、ScanNet++、HiRoom)上,均实现了位姿与几何估计的一致提升。
英文摘要
Recent Vision Foundation Models (VFMs) predict depth, camera pose, and pointmap in a single forward pass without per-scene optimization, achieving strong generalization. However, enforcing explicit multi-view geometric consistency, e.g., through bundle adjustment, is computationally costly and is thus not imposed during VFM pretraining, so such inconsistency can arise. To address this, implicit self-consistency derived from model outputs (e.g., pointmaps, features), though enforced at test-time in prior work, delivers inherently limited performance gain, especially on scenes where the pretrained VFM is highly inaccurate. In contrast to this implicit signal, we propose Self-Geometry, a plug-and-play test-time adaptation pipeline that directly imposes explicit multi-view geometric constraints using 2D pixel correspondences as pseudo ground-truth. Our proposed Self-Geometry consists of Geometric Disentanglement Optimization, which combines Multi-View Consistency and Epipolar Consistency losses with Gradient Disentanglement to prevent gradient conflict; Frame Angular-Neighbor, a view sampler based on SO(3) geodesic distances for lightly imposing these constraints; and Lightweight TTA, which adapts VFMs via LoRA. Our method achieves consistent improvements in both pose and geometry estimation across six VFMs (VGGT, $π^3$, DA3-Giant/Large/Base/Small) and four benchmarks (7Scenes, ETH3D, ScanNet++, HiRoom).
CommentsProject page: https://cmlab-korea.github.io/Self-Geometry/