发表机构
The University of Tokyo; University of Science and Technology of China; University of Zaragoza; Tohoku University(东京大学; 中国科学技术大学; 萨拉戈萨大学; 东北大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
F2SLAM将前馈几何直接转化为持久稠密因子图中的优化原生测量,通过高低频双流约束位姿与逆深度,在标定和未标定下提升轨迹精度,未标定ATE RMSE降至0.002米。
AI 中文摘要
前馈3D模型提供了强大的多视图几何先验,而在线同步定位与建图(SLAM)主要依赖局部测量,在长序列中可能累积漂移。现有结合两者的尝试通常将前馈预测视为外部几何状态,在事后与在线估计对齐或融合,这使得更广泛的多视图证据游离于优化器之外,而该优化器用于精化SLAM状态。我们提出F2SLAM,它直接将前馈几何转化为附加到持久稠密因子图上的优化原生目标权重测量。高频流维持局部跟踪约束和图连通性,而低频流利用更广泛的多视图上下文,在状态一致性检查后选择性地刷新现有测量。两条流通过单个稠密束调整约束相同的位姿、逆深度和可选相机内参。在多个基准上的实验表明,在标定和未标定设置下,轨迹估计一致地强健,稠密重建得到改进。值得注意的是,未标定配置在Replica数据集上将平均ATE RMSE从最强前馈基线的0.030米降至0.002米。
英文摘要
Feed-forward 3D models provide strong multi-view geometric priors, while on- line simultaneous localization and mapping (SLAM) relies mainly on local mea- surements and can accumulate drift over long sequences. Existing attempts to combine the two typically treat feed-forward predictions as an external geomet- ric state that is aligned or fused with the online estimate after the fact, which keeps broader multi-view evidence outside the optimizer that refines the SLAM state. We present F2SLAM, which instead converts feed-forward geometry di- rectly into optimization-native target-weight measurements attached to a persis- tent dense factor graph. A high-frequency stream maintains local tracking con- straints and graph connectivity, while a low-frequency stream uses wider multi- view context to selectively refresh existing measurements after a state-consistency check. Both streams constrain the same poses, inverse depths, and optional cam- era intrinsics through a single dense bundle adjustment. Experiments on multiple benchmarks demonstrate consistently strong trajectory estimation and improved dense reconstruction in both calibrated and uncalibrated settings. Notably, the uncalibrated configuration reduces the average ATE RMSE from 0.030 m for the strongest feed-forward baseline to 0.002 m on the Replica dataset.
Comments21pages