arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从含噪视图图学习全局相机位姿用于运动恢复结构

Learning Global Camera Poses from Noisy View-Graphs for Structure from Motion

Fadi Khatib, Meirav Galun, Ronen Basri

arXiv 2609.09491首次发表:更新:

发表机构

Weizmann Institute of Science(魏茨曼科学研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种基于图神经网络的全局运动恢复结构框架,从含噪视图图学习相机位姿,无需真值监督,高效可扩展,在多个数据集上精度优于深度轨迹法,速度优于经典流程。

AI 中文摘要

相机位姿估计是三维重建和视图合成流程中的关键步骤。我们提出了一种基于学习视图图聚合的深度全局运动恢复结构框架。我们的方法采用一个置换等变、边条件图神经网络,以含噪的相对位姿对作为输入,输出全局一致的相机外参。该网络在无真值监督的情况下训练,仅依赖相对位姿一致性目标。随后进行三维点三角化和鲁棒光束法平差。我们的方法高效,可扩展至超过一千张图像,并且对图密度具有鲁棒性。我们在MegaDepth、1DSfM、Strecha和BlendedMVS数据集上评估了我们的方法。这些实验表明,与深度基于轨迹的方法相比,我们的方法在旋转和平移精度上更优,同时在许多场景中注册了更多图像,并且与最先进的经典流程相比,取得了有竞争力的结果,同时速度更快。

英文摘要

Camera pose estimation is a key step in 3D reconstruction and view-synthesis pipelines. We present a deep, global Structure-from-Motion framework based on learned view-graph aggregation. Our method employs a permutation-equivariant, edge-conditioned graph neural network that takes noisy pairwise relative poses as input and outputs globally consistent camera extrinsics. The network is trained without ground-truth supervision, relying solely on a relative-pose consistency objective. This is followed by 3D point triangulation and robust bundle adjustment. Our approach is efficient, scalable to more than a thousand images, and robust to graph density. We evaluate our method on MegaDepth, 1DSfM, Strecha, and BlendedMVS. These experiments demonstrate that our method achieves superior rotation and translation accuracy compared to deep track-centric methods while registering more images across many scenes, and competitive results compared to state-of-the-art classical pipelines, while being much faster.

CommentsAccepted to ECCV 2026. Project page: https://vgpa-sfm.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑