FounRef:基于稀疏锚点的冻结单目基础先验的鲁棒、结构保持且快速的度量细化
FounRef: Robust, Structure-Preserving, and Fast Metric Refinement of Frozen Monocular Foundation Priors with Sparse Anchors
浏览论文内容
中文总结 AI 辅助
FounRef提出一种无需训练的模块化方法,利用稀疏锚点细化冻结单目基础先验,生成稠密度量深度,在域外数据上显著降低误差并大幅提升推理速度。
中文摘要 AI 辅助
来自相机的稠密度量深度对于现实世界的3D应用至关重要,然而同时实现准确性、忠实的表面几何和快速推理仍然具有挑战性。单目基础模型提供了丰富、可迁移的几何先验,但缺乏可靠的度量尺度,而深度补全网络以几何保真度、跨域鲁棒性或速度为代价恢复度量深度。我们提出FounRef,一种无需训练的方法,将冻结的单目基础先验与稀疏度量锚点对齐,以产生稠密度量深度。FounRef在设计上是模块化的:其深度先验、锚点来源和细化求解器均可独立替换。我们使用MoGe-2和LiDAR锚点实例化FounRef。FounRef将每个锚点与先验的稠密深度预测进行验证,拒绝由跨传感器错位引起的不一致,这些不一致是仅基于几何的滤波器无法检测到的。然后,它通过结构保持求解器应用全局和局部度量校正,保留先验的细粒度几何。FounRef无需任务特定训练,即可在陌生相机和场景中开箱即用。在域外数据上,与最先进的深度补全网络DMD3C相比,它实现了高达24%的深度误差降低、92%的表面法线噪声降低以及近15倍的推理加速。通过将度量对齐与几何预测解耦,FounRef为稠密度量深度提供了一种准确、几何忠实且高效的方法,可直接受益于基础模型和度量传感器的未来进展。
英文摘要
Dense metric depth from cameras is essential to real-world 3D applications, yet achieving accuracy, faithful surface geometry, and fast inference simultaneously remains challenging. Monocular foundation models provide rich, transferable geometric priors but lack reliable metric scale, while depth-completion networks recover metric depth at the cost of geometric fidelity, cross-domain robustness, or speed. We present FounRef, a training-free method that aligns a frozen monocular foundation prior with sparse metric anchors to produce dense metric depth. FounRef is modular by design: its depth prior, anchor source, and refinement solver can each be replaced independently. We instantiate FounRef with MoGe-2 and LiDAR anchors. FounRef validates each anchor against the prior's dense depth prediction, rejecting inconsistencies caused by cross-sensor misalignment that geometry-only filters cannot detect. It then applies global and local metric corrections through a structure-preserving solver, retaining the prior's fine-grained geometry. FounRef requires no task-specific training and operates out of the box across unfamiliar cameras and scenes. On out-of-domain data, it delivers up to 24% lower depth error, 92% lower surface-normal noise, and almost 15x faster inference than DMD3C, a state-of-the-art depth-completion network. By decoupling metric alignment from geometry prediction, FounRef provides an accurate, geometrically faithful, and efficient approach to dense metric depth that can directly benefit from future advances in foundation models and metric sensors.
发表机构
- University of the Bundeswehr Munich(慕尼黑联邦国防军大学)
机构由 AI 辅助整理,请以论文原文为准。