带有任意先验的Depth Anything
Depth Anything with Any Prior
- Zhejiang University(浙江大学)
- The University of Hong Kong(香港大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出Prior Depth Anything框架,通过从粗到精的流水线融合度量深度先验与单目深度预测结果,在7个真实数据集的多项深度任务上实现了优异的零样本泛化,还支持灵活的准确率-效率权衡。
AI中文摘要:
本研究提出Prior Depth Anything框架,该框架将深度测量中不完整但精确的度量信息,与深度预测中相对但完整的几何结构相结合,可为任意场景生成准确、稠密且细节丰富的度量深度图。为此,我们设计了一套从粗到精的流水线,以逐步整合这两种互补的深度信息源。首先,我们引入像素级度量对齐与距离感知加权机制,通过显式利用深度预测来预填充各类度量先验,有效缩小了不同先验模式间的域差距,提升了在不同场景下的泛化能力。其次,我们开发了一个条件式单目深度估计(MDE)模型,用于修正深度先验中固有的噪声。通过以归一化后的预填充先验和预测结果作为条件,该模型进一步隐式融合了这两种互补的深度信息源。我们的模型在7个真实世界数据集上,于深度补全、超分辨率和修复任务中展现出出色的零样本泛化能力,性能持平甚至超越了以往的任务专用方法。更重要的是,它在具有挑战性的未见过的混合先验上表现良好,还能通过切换预测模型实现测试时性能提升,在随MDE模型发展演进的同时,提供了灵活的准确率-效率权衡。
英文摘要:
This work presents Prior Depth Anything, a framework that combines incomplete but precise metric information in depth measurement with relative but complete geometric structures in depth prediction, generating accurate, dense, and detailed metric depth maps for any scene. To this end, we design a coarse-to-fine pipeline to progressively integrate the two complementary depth sources. First, we introduce pixel-level metric alignment and distance-aware weighting to pre-fill diverse metric priors by explicitly using depth prediction. It effectively narrows the domain gap between prior patterns, enhancing generalization across varying scenarios. Second, we develop a conditioned monocular depth estimation (MDE) model to refine the inherent noise of depth priors. By conditioning on the normalized pre-filled prior and prediction, the model further implicitly merges the two complementary depth sources. Our model showcases impressive zero-shot generalization across depth completion, super-resolution, and inpainting over 7 real-world datasets, matching or even surpassing previous task-specific methods. More importantly, it performs well on challenging, unseen mixed priors and enables test-time improvements by switching prediction models, providing a flexible accuracy-efficiency trade-off while evolving with advancements in MDE models.