单目图像的单目深度估计:进展与机遇
Monocular Depth Estimation from a Single Image: Progress and Opportunities
浏览论文内容
中文总结 AI 辅助
本综述梳理单目深度估计从早期学习方法到基础模型的演进,介绍相关数据集、方法进展,对比代表性模型,探讨其应用与未来方向。
中文摘要 AI 辅助
单目深度估计长期以来一直是计算机视觉领域的核心挑战,它能实现3D重建、机器人技术、自动驾驶和增强现实等广泛应用。本综述追溯了该领域从早期基于学习的方法到变革性基础模型出现的演进历程。我们首先明确问题,区分相对深度估计与度量深度估计,并梳理塑造了十年研究的关键挑战。接着介绍常见的问题设定及最常用的数据集,涵盖室内、室外和合成数据。随后回顾基础模型时代之前的主要进展,提炼出那些推动精度、效率和鲁棒性提升的有影响力方法的核心见解。之后转向近期基于基础模型方法的热潮,将其分为判别式和生成式范式,并强调大规模预训练(如DINOv3)和合成数据的关键作用。我们通过定量基准和定性示例对比代表性模型,探讨其向基于视频的深度估计的自然扩展。此外,为说明实际应用价值,我们重点介绍深度估计在视觉SLAM、内容生成和机器人感知等应用中的集成情况。最后,概述该领域进入基础模型时代后面临的开放挑战和有前景的研究方向。
英文摘要
Monocular depth estimation has long stood as a fundamental challenge in computer vision, enabling a wide range of applications including 3D reconstruction, robotics, autonomous driving, and augmented reality. This survey traces the field's evolution from early learning-based methods to the emergence of transformative foundation models. We begin by framing the problem, distinguishing between relative and metric depth estimation, and highlighting the key challenges that have shaped a decade of research. We then present common problem formulations and introduce the most widely used datasets, covering indoor, outdoor, and synthetic data. Following this, we review major advances prior to the foundation model era, distilling core insights from influential methods that contributed to improvements in accuracy, efficiency, and robustness. The survey then turns to the recent surge of foundation-model-based approaches, categorizing them into discriminative and generative paradigms and emphasizing the critical roles of large-scale pretraining (e.g., DINOv3) and synthetic data. We compare representative models using both quantitative benchmarks and qualitative examples, and discuss natural extensions to video-based depth estimation. Further, to illustrate real-world impact, we highlight the integration of depth estimation into applications such as visual SLAM, content generation, and robot perception. Finally, we outline open challenges and promising research directions as the field advances further into the era of foundation models.