VidMap:利用时序结构实现基于视频的运动恢复结构
VidMap: Exploiting Temporal Structure for Video-Based Structure-from-Motion
查看机构详情
- ETH Zurich(苏黎世联邦理工学院)
- Google(谷歌)
- Microsoft(微软)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
VidMap结合SLAM序列约束与SfM全局优化,利用宽基线匹配和深度先验,在多样挑战性数据集上,较最新SLAM和SfM更鲁棒准确,可实现任意长未标定视频的度量重建。
中文摘要 AI 辅助
为任意无约束视频精确恢复相机的标定参数与度量位姿,可为导航和场景理解解锁大规模训练数据。该问题的主流方法存在严重局限:同时定位与建图(SLAM)因因果增量式特性,对初始化和瞬时故障敏感,常过度优化以适配实时运行,且通常需要已知相机标定参数;而运动恢复结构(SfM)通常不考虑图像顺序,可实现最优初始化与全局优化,但对视觉对称性和极端运动缺乏鲁棒性。为弥合这一差距,本文提出一种结合SLAM强序列约束与离线SfM灵活性及全局优化的系统,可对任意长时长、未标定视频实现度量重建。该系统利用宽基线密集图像匹配的最新进展,将时序排序作为可靠回环检测的核心要素,并通过度量单目深度先验增强全局优化。对包含极端运动和视觉对称性的多样挑战性数据集的全面评估表明,本文方法在给定或未知相机标定参数的情况下,相较于经典或学习型的最新SLAM和SfM方法,鲁棒性和准确性显著更优。代码已公开于https URL。
英文摘要
Accurately recovering the camera's calibration and metric poses for any unconstrained video would unlock large-scale training data for navigation and scene understanding. The dominant approaches to this problem are severely limited: Simultaneous Localization and Mapping (SLAM) is sensitive to initialization and transient failures due to its causal, incremental nature; it is often over-optimized for real-time operation and generally requires known camera calibration; while Structure-from-Motion (SfM) typically forgoes any image ordering, enabling optimal initialization and global optimization, but lacks robustness to visual symmetries and extreme motions. To bridge this gap, we introduce a system that combines the strong sequential constraints of SLAM with the flexibility and global optimization of offline SfM, enabling the metric reconstruction of arbitrary, long, uncalibrated videos. This system leverages recent advances in wide-baseline dense image matching, treats temporal ordering as a first-class citizen for reliable loop closure, and augments global optimization with metric monocular depth priors. As a result, thorough evaluations on diverse, challenging datasets that exhibit extreme motion and visual symmetries reveal that our approach is significantly more robust and accurate than both state-of-the-art SLAM and SfM, classical or learned, with given or unknown camera calibration. The code is publicly available at https://github.com/cvg/vidmap.