arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2401.10891cs.CV

Depth Anything:释放大规模无标签数据的力量

Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data

  • HKU(香港大学)
  • TikTok(字节跳动旗下TikTok)
  • CUHK(香港中文大学)
  • ZJU(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

Lihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu, Jiashi Feng, Hengshuang Zhao

更新

AI总结:

本文提出 Depth Anything,通过数据引擎自动标注约 6200 万无标签数据,并结合数据增强与语义先验辅助监督,构建出具备强零样本泛化能力的单目深度估计基础模型,在多个数据集上取得 SOTA。

AI中文摘要:

本文提出了 Depth Anything,一种用于鲁棒单目深度估计的高度实用解决方案。我们不追求新颖的技术模块,而是旨在构建一个简单而强大的基础模型,以处理任何情况下的任何图像。为此,我们通过设计一个数据引擎来收集并自动标注大规模无标签数据(约 6200 万),从而扩大数据集规模,显著增大了数据覆盖范围,因此能够降低泛化误差。我们研究了两种简单而有效的策略,使数据规模化扩展具有前景。首先,通过利用数据增强工具创建了一个更具挑战性的优化目标。这迫使模型主动寻找额外的视觉知识并获得鲁棒表示。其次,开发了一种辅助监督,以强制模型从预训练编码器中继承丰富的语义先验。我们广泛评估了其零样本能力,包括六个公共数据集和随机拍摄的照片。它展示了令人印象深刻的泛化能力。此外,通过使用来自 NYUv2 和 KITTI 的度量深度信息对其进行微调,建立了新的 SOTA。我们更好的深度模型也带来了更好的深度条件 ControlNet。我们的模型已在 https://github.com/LiheYoung/Depth-Anything 发布。

英文摘要:

This work presents Depth Anything, a highly practical solution for robust monocular depth estimation. Without pursuing novel technical modules, we aim to build a simple yet powerful foundation model dealing with any images under any circumstances. To this end, we scale up the dataset by designing a data engine to collect and automatically annotate large-scale unlabeled data (~62M), which significantly enlarges the data coverage and thus is able to reduce the generalization error. We investigate two simple yet effective strategies that make data scaling-up promising. First, a more challenging optimization target is created by leveraging data augmentation tools. It compels the model to actively seek extra visual knowledge and acquire robust representations. Second, an auxiliary supervision is developed to enforce the model to inherit rich semantic priors from pre-trained encoders. We evaluate its zero-shot capabilities extensively, including six public datasets and randomly captured photos. It demonstrates impressive generalization ability. Further, through fine-tuning it with metric depth information from NYUv2 and KITTI, new SOTAs are set. Our better depth model also results in a better depth-conditioned ControlNet. Our models are released at https://github.com/LiheYoung/Depth-Anything.

补充信息

↑