arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.08084cs.CVcs.LG

Marigold V2:重新审视扩散变换器用于单目深度估计

Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation

  • EPFL(洛桑联邦理工学院)
  • HUAWEI Bayer Lab(华为拜耳实验室)
  • University of Bologna(博洛尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

Igor Pavlovic, Thiemo Wandel, Anton Obukhov, Luca Bartolomei, Andrey Davydov, Fabio Tosi, Matteo Poggi, Sabine Süsstrunk, Dengxin Dai

中文总结 AI 辅助

Marigold V2利用扩散变换器,通过语义对齐和Sinkhorn损失的两阶段微调,实现单步推理,在KITTI和ETH3D上AbsRel改善16-26%,并提升深度、法线和内在分解的精度。

中文摘要 AI 辅助

单目深度估计是一项普遍存在但高度不适定的计算机视觉任务,其下游应用包括场景重建、计算摄影和机器人等领域。尽管该领域已相当成熟,但最近的模型仍难以泛化到分布外输入,并难以生成清晰且细节丰富的深度图。在本文中,我们重新审视了Marigold,这是一套利用扩散变换器(DiT)架构,将现代图像生成和编辑模型改造为最先进的单目深度估计器的技术。我们的方案针对从预训练的多步流匹配模型进行单步推理,并在需要时进行量化,从而在保持模型能力的同时,运行成本低廉。我们分析了朴素训练产生的伪影,并确定了两种有效的补救措施:将模型的内部表示与从真实数据中提取的语义特征对齐,以及采用围绕新型基于Sinkhorn的损失构建的两阶段微调协议。结果是更清晰、更干净的深度图,能够很好地泛化到分布外数据,在KITTI和ETH3D上,与之前的最佳结果相比,AbsRel改善了16-26%。在定性方面,我们的模型解决了先前模型无法处理的皮毛、树叶和细发边缘。此外,Marigold V2在应用于其他密集回归任务(如表面法线估计和内在图像分解)时,也取得了最先进的结果。项目网站:此https URL

英文摘要

Monocular depth estimation is a ubiquitous yet highly ill-posed computer vision task, with downstream applications in scene reconstruction, computational photography, and robotics, among others. Despite the field's maturity, recent models still struggle to generalize to out-of-distribution inputs and to produce sharp and detailed depth maps. In this paper, we revisit Marigold, a set of techniques for repurposing modern image generation and editing models, powered by the diffusion transformer (DiT) architecture, into state-of-the-art monocular depth estimators. Our recipes target single-step inference from pretrained multi-step flow-matching models, with quantization where needed, preserving model capacity while remaining cheap to run. We analyze the artifacts of naive training and identify two effective remedies: aligning the model's internal representations with semantic features extracted from ground-truth, and adopting a 2-stage fine-tuning protocol built around a novel Sinkhorn-based loss. The results are crisper, cleaner depth maps that generalize well out-of-distribution, with 16-26% improvement in AbsRel over the previous best on KITTI and ETH3D. Qualitatively, our model resolves fur, foliage, and hair-thin edges that have eluded prior models. Furthermore, Marigold V2 achieves state-of-the-art results when applied to other dense regression tasks, such as surface normals estimation and intrinsic image decomposition. Project website: https://hf.co/spaces/huawei-bayerlab/marigold-v2-web

补充信息

相关深度报道

↑