arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.04554cs.CV

EagleDepth:通过像素扩散解码器实现高效细粒度深度估计

EagleDepth: Efficient Fine-Grained Depth Estimation via Pixel Diffusion Decoder

Bowen Chai, Tianbao Zhang, Shuyu Wu, Dexin Zuo, Zhaoxin Fan, Danping Zou

首次发表
浏览论文内容

中文总结 AI 辅助

EagleDepth提出结合潜在扩散几何先验与像素扩散解码器的高效高分辨率单目深度估计框架,在多个数据集上达到最先进性能,推理更快且保留细节。

中文摘要 AI 辅助

从高分辨率图像中恢复详细的几何结构对于精确感知周围环境和物体至关重要。然而,现有方法使用潜在空间建模和VAE重建可能会损害几何细节。此外,从潜在编码解码会引入大量的推理开销。为了解决这些问题,我们提出了EagleDepth,一个用于高分辨率单目深度估计的高效框架,它结合了潜在扩散的几何先验与细粒度像素空间生成。我们的关键思想是保留深度感知的潜在表示作为指导,同时在像素空间中直接生成最终深度图。我们按顺序训练潜在和像素组件:首先,使用配对的RGB-深度监督微调预训练的潜在扩散模型;然后,调整预训练的像素扩散解码器PiD,以基于学习到的特征预测深度。像素组件的训练从1024分辨率开始,并持续在多个分辨率上进行,最高可达4K。潜在分支处理调整大小的低分辨率RGB图像,而像素分支在目标分辨率下生成深度,绕过原始的VAE解码器。这种设计保留了学习到的几何知识,而无需潜在骨干在输出分辨率下运行。在五个常用的深度估计数据集和高分辨率Synth4K数据集上,我们的框架实现了最先进的深度估计性能,具有更快的推理速度和更好的细结构及物体边界保留。

英文摘要

Recovering detailed geometry from high-resolution images is critical for precise perception of the surroundings and objects. However, existing methods which use latent-space modeling and VAE reconstruction can compromise geometric details. Furthermore, decoding from latent codes introduces substantial inference overhead. To address those issues, we present EagleDepth, an efficient framework for high-resolution monocular depth estimation that combines the geometric priors of latent diffusion with fine-grained pixel-space generation. Our key idea is to retain depth-aware latent representations as guidance while generating the final depth map directly in pixel space. We train the latent and pixel components sequentially: first, we fine-tune a pretrained latent diffusion model using paired RGB--depth supervision; then, we adapt a pretrained pixel diffusion decoder, PiD, to predict depth conditioned on the learned features. Training of the pixel component starts at 1024 resolution and continues across multiple resolutions up to 4K. The latent branch processes resized, lower-resolution RGB images, while the pixel branch generates depth at the target resolution, bypassing the original VAE decoder. This design preserves learned geometric knowledge without requiring the latent backbone to operate at the output resolution. On five commonly used depth estimation datasets and the high-resolution Synth4K dataset, our framework achieves state-of-the-art depth estimation performance, with faster inference and better preservation of fine structures and object boundaries.

发表机构

  • Shanghai Jiao Tong University(上海交通大学)
  • Beihang University(北京航空航天大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑