发表机构
ITMO University(ITMO大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DAVIO利用单一多视图深度模型实现稠密单目惯性SLAM,通过无特征线性系统快速启动VIO滤波器,并用度量姿态条件化深度模型建图,在EuRoC和ORI序列上显著优于现有方法。
AI 中文摘要
相机和IMU是进行度量定位和稠密建图的最小传感器配置,然而经典的视觉-惯性滤波器必须等待视差出现才能启动,并且只能保留稀疏地标。相比之下,前馈几何模型可以从少量图像预测稠密结构,但既不提供度量尺度也不提供重力。我们提出了DAVIO,它使用一个单一的多视图深度模型——Depth Anything 3——同时用于启动和建图。在启动阶段,一个五图像窗口和预积分的IMU测量形成一个无特征的线性系统。其鲁棒的、经过条件数检查的解通过缓冲重放引导VIO滤波器启动。在跟踪过程中,滤波器的度量姿态对深度模型进行条件约束。残差尺度仅沿观测射线进行校正,这保持了度量相机基线,同时一个保持重力的子地图图结构,带有漂移门控的重访,对地图进行细化。在EuRoC数据集上,DAVIO启动明显更早,降低了定位误差,并且在给定相同姿态的情况下,比最先进的前馈建图器建图更准确。在建筑规模的ORI序列上,DAVIO在相同里程计下与最先进的建图器相当或更好,并且当真实姿态被实际里程计替代时,性能下降远小于它们。我们将DAVIO的代码(一个实时的稠密度量SLAM系统)发布给社区。
英文摘要
A camera and an IMU are the minimal sensor setup for metric localization and dense mapping, yet classical visual--inertial filters must wait for parallax before they start and then retain only sparse landmarks. Feed-forward geometry models, in contrast, predict dense structure from a few images but provide neither metric scale nor gravity. We present DAVIO, which uses a single multi-view depth model, Depth Anything~3, for both start-up and mapping. At start-up, a five-image window and preintegrated IMU measurements form a feature-free linear system. Its robust, conditioning-checked solution bootstraps a VIO filter through buffered replay. During tracking, the filter's metric poses condition the depth model. Residual scale is corrected only along viewing rays, which preserves the metric camera baselines, and a gravity-preserving submap graph with drift-gated revisits refines the map. On EuRoC, DAVIO starts markedly earlier, reduces the localization error, and maps more accurately than SOTA feed-forward mappers given identical poses. On building-scale ORI sequences, DAVIO is on bar or better than SOTA mappers on the same odometry, and degrades far less when GT poses are replaced by real odometry. We release the code of DAVIO, a real-time dense metric SLAM system, to the community.