发表机构
Institute of Mathematics of the Romanian Academy; POLITEHNICA Bucharest(罗马尼亚科学院数学研究所; 布加勒斯特理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出一种仅依赖点间线性关系、无需相机参数的多视角3D重建方法,通过鲁棒估计自回归矩阵,利用噪声单目深度或多帧2D投影,高效恢复3D结构。
AI 中文摘要
我们提出了一种高效且鲁棒的3D几何重建方法,该方法仅基于给定点集之间与相机无关的线性关系,这些关系随时间稳定,并通过多个点匹配进行鲁棒估计。我们本质上从多帧图像中的点对应关系中学习一个线性几何自回归矩阵$\mathbf{W}$,该矩阵确立了3D空间中一个点如何表示为所有其他点的线性组合。该矩阵是常数,不依赖于世界坐标系或相机姿态——它是点集的内在属性。我们还证明了$\mathbf{W}$的主特征向量(其特征值均为1)提供了3D点配置的齐次表示。我们方法的第一版本利用噪声单目深度图,从多帧中鲁棒地估计3D点之间线性关系的几何自回归矩阵$\mathbf{W}$。因此,我们基于深度学习的最新进展,这些进展现在提供了快速但通常有噪声的单目深度估计模型。我们的方法通过多帧上的鲁棒线性估计来处理噪声。我们方法的第二版本不需要单目深度估计图。它适用于弱透视投影的情况,即3D点之间的线性组合可以从它们在图像中的2D投影中鲁棒地估计。注意,相机投影矩阵在我们的推导中从未使用。因此,我们的方法不恢复相机姿态,而仅恢复3D结构。这是我们方法与3D几何重建相关文献之间的关键区别。
英文摘要
We present an efficient and robust method for 3D geometric reconstruction that is based solely on the camera-independent linear relationships among a given set of points, which are stable over time and robustly estimated using multiple point matches. We essentially learn, from correspondences between points across several frames, a linear geometric auto-regression matrix $\mathbf{W}$, which establishes how a point in 3D can be expressed as a linear combination of all the others. This matrix is constant and does not depend on the world coordinate system or the camera pose---it is an intrinsic property of the point set. We also show that the principal eigenvectors of $\mathbf{W}$, which all have eigenvalue $1$, provide a homogeneous representation of the 3D point configuration. The first version of our method takes advantage of noisy monocular depth maps in order to obtain, from multiple frames, a robust geometric auto-regression matrix $\mathbf{W}$ of linear relationships between the 3D points. Thus, we build on recent advances in deep learning, which now provide monocular depth estimation models that are fast but very often noisy. Our approach handles noise through robust linear estimation over several frames. The second version of our method does not need monocular depth estimation maps. It applies in cases of weak-perspective projection, when the linear combinations between the 3D points can be robustly estimated from their 2D projections in the image. Note that the camera projection matrix is never used in our derivations. Consequently, our method does not recover camera pose, but only 3D structure. This is a key difference between our method and the related literature on 3D geometric reconstruction.
Comments12 pages