发表机构
IBISC Laboratory, Université d’Évry Paris-Saclay, Université Paris-Saclay(IBISC实验室,埃夫里-巴黎萨克雷大学,巴黎萨克雷大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对单目基础模型缺乏稳定度量尺度的问题,提出VIDAR框架,结合SVO+IMU里程计与深度任意模型3,利用视觉惯性前端作度量锚点,经姿态注入等方法,在EuRoC和TUM RGB-D数据集上实现度量密集单目重建。
AI 中文摘要
单目基础模型可提供密集几何信息,但通常缺乏稳定的度量尺度。本文提出了VIDAR,这是一个视觉惯性密集重建框架,它将SVO+IMU里程计与深度任意模型3相结合。VIDAR使用视觉惯性前端作为度量锚点,提供相机姿态、尺度和一致的世界框架,以跨时间对齐密集基础模型预测。基础模型则贡献详细的局部几何信息,融合到全局重建中。研究了姿态条件下的DA3和解耦对齐策略。在EuRoC上,姿态注入将尺度误差降低到约1%,平均F@0.10达到0.463;解耦混合方法在没有地面真值姿态的情况下将其提高到0.676。在EuRoC和TUM RGB-D上的结果表明,VIDAR是实现度量密集单目重建的实用途径。
英文摘要
Monocular foundation models provide dense geometry but usually lack a stable metric scale. This paper presents VIDAR, a visual-inertial dense reconstruction framework that couples SVO+IMU odometry with Depth Anything 3. VIDAR uses the visual-inertial front end as a metric anchor: it provides camera poses, scale, and a consistent world frame for aligning dense foundation-model predictions across time. The foundation model then contributes detailed local geometry that is fused into a global reconstruction. We study both pose-conditioned DA3 and a decoupled alignment strategy. On EuRoC, pose injection reduces scale error to about 1\% and reaches 0.463 mean F@0.10; the decoupled hybrid improves this to 0.676 without ground-truth poses. Results on EuRoC and TUM RGB-D show that VIDAR is a practical route to metric dense monocular reconstruction.