发表机构
Shanghai Jiao Tong University(上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
NaCR利用NeRF与相机射线回归的互补对偶性,通过三种增强、新视角数据增强和闭环可微监督,提升视觉定位精度,在室内外基准上达到有竞争力性能。
AI 中文摘要
视觉定位(VL)是虚拟现实等视觉应用的基础技术。最近,一种新颖的视觉定位范式——相机射线回归(CRR)应运而生,它将2D图像块映射到3D相机射线,但其精度有限。为了提高CRR的精度,我们注意到一个引人注目的对偶性:该映射的逆过程本质上由新视角合成模型,即神经辐射场(NeRF)执行。NeRF通过可微射线行进从相机射线渲染图像块,而CRR则从图像块预测射线。受这种互补关系的启发,我们提出了NeRF辅助相机射线回归(NaCR),一个在射线层面无缝桥接NeRF和CRR的统一框架。首先,NaCR将三种简单而有效的增强融入CRR基线。其次,利用预训练的NeRF,NaCR通过合成针对高效补丁级消费定制的新视角来增强训练数据。最后,利用NeRF的可微性,NaCR形成了一个闭环监督流水线,其中光度渲染误差被反向传播以优化预测的相机射线。为了在高度非凸的图像空间中确保稳定收敛,我们引入了一个两阶段训练课程。跨室内和室外基准的大量实验表明,NaCR达到了有竞争力的精度。全面的消融研究验证了每个提出组件的有效性。
英文摘要
Visual localization (VL) is a fundamental technology for vision applications such as virtual reality. Recently, a novel VL paradigm, Camera Ray Regression (CRR), has emerged, which maps 2D image patches to 3D camera rays, but its accuracy is limited. To improve CRR accuracy, we notice a compelling duality: the inverse of this mapping is inherently performed by the novel view synthesis model, \ie, Neural Radiance Fields (NeRF). While NeRF renders image patches from camera rays via differentiable ray marching, CRR predicts the rays from image patches. Motivated by this complementary relationship, we propose NeRF-aided Camera Ray Regression (NaCR), a unified framework that seamlessly bridges NeRF and CRR at the ray level. First, NaCR incorporates three simple yet effective enhancements into the CRR baseline. Second, leveraging a pre-trained NeRF, NaCR augments the training data by synthesizing novel views tailored for efficient, patch-level consumption. Finally, exploiting the differentiability of NeRF, NaCR forms a closed-loop supervision pipeline where photometric rendering errors are back-propagated to optimize the predicted camera rays. To ensure stable convergence within the highly non-convex image space, we introduce a two-stage training curriculum. Extensive experiments across indoor and outdoor benchmarks demonstrate that NaCR achieves competitive accuracy. Comprehensive ablation studies validate the efficacy of each proposed component.
Commentsv0